[root@02 ~]# dmraid -s *** Group superset .ddf1_disks --> Subset name : ddf1_4c5349202020202010000411100010044711471181ddbcc4 size : 123045888 stride : 128 type : stripe status : ok subsets: 0 devs : 1 spares : 0
Tried to deactivate the raid device
[root@02 ~]# dmraid -an ERROR: dos: partition address past end of RAID device The dynamic shared library "libdmraid-events-ddf1.so" could not be loaded: libdmraid-events-ddf1.so: cannot open shared object file: No such file or directory
Remove all raid devices metadata
[root@02 ~]# dmraid -r -E /dev/sdb Do you really want to erase "ddf1" ondisk metadata on /dev/sdb ? [y/n] :y ERROR: ddf1: seeking device "/dev/sdb" to 32779907366912 ERROR: writing metadata to /dev/sdb, offset 64023256576 sectors, size 0 bytes returned 0 ERROR: erasing ondisk metadata on /dev/sdb
Format the entire disk with zeros so that it does not get detected as raid device and cause further problem.
As the number of metrics in your environment grows you start to see huge impact in system IO performance. At one stage the disk utilization stays at 100% and cpu spends a lot of time waiting for IO. At this time the IO exceeds the theoratical IO supported by the disk. We might go ahead and add fast 15K disks or raid10 array. The disk contention stays as long as we keep on expanding the cluster and adding new metrics.
A simple solution would be to add the rrdcached layer in the middle. There are certain things to consider such as updating the rrdtool package and recompiling ganglia with the new rrdtool support.
Just a step by step guide of doing it.
A. Steps for building and install rrdtool and ganglia.
1. Since gmetad runs with ganglia user and rrdcached require access to write to rrd dir and apache needs access of same directory. Add ganglia to apache group.
9. Edit td-agent configuration in log aggregator server
<source>
type forward
port 24224
</source>
<match stats.access>
type webhdfs
host <NAMENODE OR HTTPFS HOST>
port 14000
path /user/hdfs/stats_logs/stats_access.%Y%m%d_%H.log
httpfs true
username httpfsuser
</match>
10. Start td-agent in log aggregator host
/etc/init.d/td-agent start
* ensure that there are no errors in /var/log/td-agent/td-agent.log
While starting namenode you might come across this error.
Class org.apache.hadoop.thriftfs.NamenodePlugin not found
In cdh4 it does not require a plug-in on the NameNode or DataNodes. Hence all the configuration related to that should be removed from namenode and datanode hdfs-site.xml
Extract data from oplog in MongoDB and restore in another MongoDB server.
Recently I came across a problem where we have to do a lot of modifications in the mongodb server which will be having issues with the production database. We removed the replication between the master and slave and then did the operation in slave and then updated the data using the tool wordnik-oss tools.
Unfortunately we did not have a replica set and we had normal Master-Slave setup. While all the updates are happening in slave I need to keep track of the master data so that I can add it to the slave. For this I used a tool named mongodb-admin-utils in wordnik-oss https://github.com/wordnik/wordnik-oss.
Required software:
1. Java and Git: yum install java-1.6.0-sun-devel java-1.6.0-sun git
2. Maven: recent version of wordnik-oss require maven 3
cd /usr/src wgethttp://apache.techartifact.com/mirror/maven/binaries/apache-maven-3.0.4-bin.tar.gz tar zxf apache-maven-3.0.4-bin.tar.gz
2. Compile and build In my case I only needed mongodb-admin-utils and hence I packaged only that.
cd wordnik/modules/mongo-admin-utils /usr/src/apache-maven-3.0.4/bin/mvn package
Once this is complete you can use mongo-admin-utils in the host.
Get Incremental oplog Backup from mongo master server
cd wordnik/modules/mongo-admin-utils ./bin/run.sh com.wordnik.system.mongodb.IncrementalBackupUtil -o /root/mongo -h mastermongodb
/root/mongo => output directory where the oplog is stored. mastermongodb => mongodb master host.
** We can't use this tool in slave as there is no oplog in slave.
Replay the Data from the oplog to the database
I had some problems in restoring data from backup and I had to add the following settings for the restore to work without any issues. ulimit -n 20000
Added the following Java options in run.sh so that it does not fail with Out Of Memory (OOM ) erros. JAVA_CONFIG_OPTIONS="-Xms5g -Xmx10g -XX:NewSize=2g -XX:MaxNewSize=2g -XX:+UseConcMarkSweepGC -XX:+UseParNewGC -XX:PermSize=2g -XX:MaxPermSize=2g"
Replay Command:
./bin/run.sh com.wordnik.system.mongodb.ReplayUtil -i /root/mongo -h localhost localhost => the mongodb server in which you want the data to be added.
If you face any issues you can go ahead and file a issue in https://github.com/wordnik/wordnik-os . The developer is a awesome person and will help you sort out the issue.
When you want to use custom hostname for puppet it shows the following error.
=============
err: Could not retrieve catalog from remote server: hostname was not match with the server certificate warning: Not using cache on failed catalog err: Could not retrieve catalog; skipping run err: Could not send report: hostname was not match with the server certificate ==============
In my case I wanted to use the default hostname "puppet" . Add the following entries to puppet master configuration file /etc/puppet/puppet.conf
I usually upgrade puppet using rpm package and always love to stay on latest stable. The EPEL repo usually does not update the repositories as I want. Recently I wanted to update to latest stable 2.7.6
While I was installing 2.7.6 I had issues that the rpmbuild failed with the following error. ================== sed: can't read lib/puppet/network/http_server/mongrel.rb: No such file or directory =================
Replacing http_server with http in line number 79 in the puppet.spec file fixed the issue.
Recently I migrated to Reliance 3G services. The installation disk did not come with any applet for linux systems. Hence I had a little trouble connecting it to the internet using default network manager. You can use rcomnet APN.
I thought of getting a applet like the one they have for windows and mac. After some searching I came across a blog post which covered it.
Download the following installation file linux.zip ( DOWNLOAD LINK )
unzip linux.zip
cd linux
chmod +x install
./install
Once it is done then it will install the drivers required for the device to be operational and also installs a applet for movistar mobile network which is widely used in europe.
The applet will automatically pop-up when you insert the datacard. Click connect. Change the language to english from the menu and you should be able to use a english interface.
In an automated environment where new instances are added automatically and manged by puppet it is a great problem when the puppet master has some issues. It can act as a SPOF.
I happened as a accidental problem that puppet master had a 100% disk usage. As a result the requests from puppet clients of new instances were failing with 503 error.
On checking the puppet master I could see the following error in puppet master error log.
=============================
Exception PhusionPassenger::UnknownError in PhusionPassenger::Rack::ApplicationSpawner (header too long (OpenSSL::X509::CRLError)) (process 598, thread #):
=============================
We have replaced passenger instead of the built in webrick for performance. Now checking the master there were no error. Accidentally when I tried to list out the certificates that are there in the host I got the following error. ============================= puppetca --list --all
err: Could not call list: header too long
=============================
Searching the forums I could see that this can happen if there were 0 byte certificate requests in /var/puppet/ssl/ca/requests or ( /var/lib/puppet/ssl/ca/requests ). In our case it was the /etc/puppet/ssl/ca/ca_crl.pem which was 0 byte. Removed the file and everything was back to normal.
It is quite a bad day when the master of automation gets involved in some kind of trouble.
We have lot of scripts which do automatic maintenance work during weekends. Eventhough the scripts are written to take care of errors it doesn't have a option to notify nagios that the maintenance work is taking place.
The person who is Oncall also gets frustrated seeing the alerts disturbing his weekend peace. He might even screw up the entire maintenance taking place.
Hence we needed the script to notify nagios that a maintenance is taking place and not to send out notifications.
We were using nagios3 as the monitoring service. The great command line utility curl came in handy here.
We use curl to send a POST request to the nagios admin interface emulating a user experience.
Sometimes the Gmond process does not start and spews the following error.
==========================
gmond -d 10
udp_recv_channel mcast_join=10.16.101.81 mcast_if=NULL port=8664 bind=10.16.101.81
Error creating multicast server mcast_join=10.16.101.81 port=8664 mcast_if=NULL family='inet4'. Exiting.
==========================
This happens due to some multicast routing issues. I am not sure exactly what is causing this problem. The fix is to explicitly add a route.
============================
route add -host 239.2.11.71 dev eth0
=============================
Need to learn what is causing this problem though..
Usually the code should be tested in a development environment before pushing the code to production. Automating test process is an important process in deployment. CCRB is a great tool to do this.
The code base is written in such a way that it can be deployed to development or production based on environment variables passed using the capistrano deployment script.
The development deployment initiates a CCRB build and testing process in the development cruisecontrol project which has the same code base. During this process the CCRB should be capable of invoking a development environment variable.
In comman setups we have the environmental variable 'development' and 'production' to differentiate the between production and development.
We add the following entries to cruise_config.rb to pass the 'development' environmental variables to the ccrb build.
============
ENV['env'] = 'development'
============
You can create a file named build_requested in project rootdir to initiate a build process.
Capistrano does not accept ruby methods. Suppose I need to get user input I can't use gets.strip and it would spew Method not found error.
You can use the following method to get the user input in capistrano deploy scripts.
==============
puts "This is a critical code do you want to proceed (y/n)"
value = STDIN.gets[0..0] rescue nil
exit unless value == 'y' or value == 'Y'
===============
Capistrano is full and fast automation solution. Don't include too much of user interaction in that unless necessary.
Just difficult for me to remember this thing. This is a way to execute remote command which can executed only using sudo privileges.
===============
ssh -t testuser@testserver "/usr/bin/sudo sh -c w"
===============
This is a guide to zero security in linux. It is applicable only in places where security is not a threat and there is no threat from external networks.
The following options would allow a system running ssh to have a user with empty password so that you can use this user to login without any password or ssh keys.
I am adding a test user for this purpose.
useradd noneknowme
passwd -d noneknowme
Configure ssh server to allow empty passwords.
Edit the following line in /etc/ssh/sshd_config
================
PermitEmptyPasswords yes
================
Restart sshd using /etc/init.d/sshd restart
Now you should be able to access the host with ssh with username noneknowme without any issues.
===============
ssh -l noneknowme test
[noneknowme@test ~]$
================
Don't try this . I am using this as a reference. Thanks to linuxquestions.org
I was wondering what is the best way to do "Zero Padding" in Bash. Suppose you need to iterate through a for loop and you need 01, 02, 03 sequence instead of 1, 2,3 then that is Zero padding.
Some of the methods I could get from the internet.