Friday, November 09, 2012

Ubuntu - Nagios Installation

This guide is intended to provide you with simple instructions on how to install Nagios from source (code) on Ubuntu and have it monitoring your local machine inside of 20 minutes. No advanced installation options are discussed here - just the basics that will work for 95% of users who want to get started.

These instructions were written based on an Ubuntu 6.10 (desktop) installation. They should work for an Ubuntu 7.10 install as well. I used a virtual machine running Ubuntu JeOS back in 2009.

If you follow these instructions, here's what you'll end up with:

Nagios and the plugins will be installed underneath /usr/local/nagios
Nagios will be configured to monitor a few aspects of your local system (CPU load, disk usage, etc.)
The Nagios web interface will be accessible at http://localhost/nagios/

Required Packages

Make sure you've installed the following packages on your Ubuntu installation before continuing.

Apache 2
GCC compiler and development libraries
GD development libraries
You can use apt-get to install these packages by running the following commands:

sudo apt-get install apache2
sudo apt-get install build-essential


With Ubuntu 6.10, install the gd2 library with this command:

sudo apt-get install libgd2-dev

With Ubuntu 7.10, the gd2 library name has changed, so you'll need to use the following:

sudo apt-get install libgd2-xpm-dev

1) Create Account Information

Become the root user.

sudo -s

Create a new nagios user account and give it a password.

/usr/sbin/useradd nagios
passwd nagios


On Ubuntu server edition (6.01 and possible newer versions), you will need to also add a nagios group (it's not created by default). You should be able to skip this step on desktop editions of Ubuntu.

/usr/sbin/groupadd nagios
/usr/sbin/usermod -G nagios nagios


Create a new nagcmd group for allowing external commands to be submitted through the web interface. Add both the nagios user and the apache user to the group.

/usr/sbin/groupadd nagcmd
/usr/sbin/usermod -G nagcmd nagios
/usr/sbin/usermod -G nagcmd www-data


2) Download Nagios and the Plugins

Create a directory for storing the downloads.

mkdir ~/downloads
cd ~/downloads


Download the source code tarballs of both Nagios and the Nagios plugins (visit http://www.nagios.org/download/ for links to the latest versions). At the time of writing, the latest versions of Nagios and the Nagios plugins were 3.0 and 1.4.11, respectively.


wget http://prdownloads.sourceforge.net/nagios/nagios-3.0.1.tar.gz
wget http://osdn.dl.sourceforge.net/sourceforge/nagiosplug/nagios-plugins-1.4.11.tar.gz



3) Compile and Install Nagios

Extract the Nagios source code tarball.

cd ~/downloads
tar xzf nagios-3.0.tar.gz
cd nagios-3.0


Run the Nagios configure script, passing the name of the group you created earlier like so:

./configure --with-command-group=nagcmd

Compile the Nagios source code.

make all


Install binaries, init script, sample config files and set permissions on the external command directory.

make install
make install-init
make install-config
make install-commandmode


Don't start Nagios yet - there's still more that needs to be done...

4) Customize Configuration

Sample configuration files have now been installed in the /usr/local/nagios/etc directory. These sample files should work fine for getting started with Nagios. You'll need to make just one change before you proceed...

Edit the /usr/local/nagios/etc/objects/contacts.cfg config file with your favorite editor and change the email address associated with the nagiosadmin contact definition to the address you'd like to use for receiving alerts.

vi /usr/local/nagios/etc/objects/contacts.cfg

5) Configure the Web Interface

Install the Nagios web config file in the Apache conf.d directory.

make install-webconf

Create a nagiosadmin account for logging into the Nagios web interface. Remember the password you assign to this account - you'll need it later.

htpasswd -c /usr/local/nagios/etc/htpasswd.users nagiosadmin

Restart Apache to make the new settings take effect.

/etc/init.d/apache2 reload

6) Compile and Install the Nagios Plugins

Extract the Nagios plugins source code tarball.

cd ~/downloads
tar xzf nagios-plugins-1.4.11.tar.gz
cd nagios-plugins-1.4.11


Compile and install the plugins.

./configure --with-nagios-user=nagios --with-nagios-group=nagios
make
make install


7) Start Nagios

Configure Nagios to automatically start when the system boots.

ln -s /etc/init.d/nagios /etc/rcS.d/S99nagios

Verify the sample Nagios configuration files.

/usr/local/nagios/bin/nagios -v /usr/local/nagios/etc/nagios.cfg

If there are no errors, start Nagios.

/etc/init.d/nagios start

8) Login to the Web Interface

You should now be able to access the Nagios web interface at the URL below. You'll be prompted for the username (nagiosadmin) and password you specified earlier.

http://localhost/nagios/

Click on the "Service Detail" navbar link to see details of what's being monitored on your local machine. It will take a few minutes for Nagios to check all the services associated with your machine, as the checks are spread out over time.

9) Other Modifications

If you want to receive email notifications for Nagios alerts, you need to install the mailx (Postfix) package.

sudo apt-get install mailx

You'll have to edit the Nagios email notification commands found in /usr/local/nagios/etc/objects/commands.cfg and change any '/bin/mail' references to '/usr/bin/mail'. Once you do that you'll need to restart Nagios to make the configuration changes live.

sudo /etc/init.d/nagios restart


Configuring email notifications is outside the scope of this documentation. Refer to your system documentation, search the web, or look to the NagiosCommunity.org wiki for specific instructions on configuring your Ubuntu system to send email messages to external addresses.

Ubuntu 101

From old Yahoo Notes I found this good intro to Ubuntu VMs.

Logs: A "live" view of a logfile on Linux

tail -f /path/thefile.log

Storage: Adding Hard drive

jblanco@us01:~$ sudo fdisk -l

Disk /dev/sda: 10.7 GB, 10737418240 bytes
255 heads, 63 sectors/track, 1305 cylinders
Units = cylinders of 16065 * 512 = 8225280 bytes

Device Boot Start End Blocks Id System
/dev/sda1 * 1 1247 10016496 83 Linux
/dev/sda2 1248 1305 465885 5 Extended
/dev/sda5 1248 1305 465853+ 82 Linux swap / Solaris

Disk /dev/sdb: 42.9 GB, 42949672960 bytes
255 heads, 63 sectors/track, 5221 cylinders
Units = cylinders of 16065 * 512 = 8225280 bytes

Disk /dev/sdb doesn't contain a valid partition table
jblanco@us01:~$ fdisk /dev/sdb

Unable to open /dev/sdb
jblanco@us01:~$ sudo fdisk /dev/sdb
Device contains neither a valid DOS partition table, nor Sun, SGI or OSF disklabel
Building a new DOS disklabel. Changes will remain in memory only,
until you decide to write them. After that, of course, the previous
content won't be recoverable.

The number of cylinders for this disk is set to 5221.
There is nothing wrong with that, but this is larger than 1024,
and could in certain setups cause problems with:
1) software that runs at boot time (e.g., old versions of LILO)
2) booting and partitioning software from other OSs
(e.g., DOS FDISK, OS/2 FDISK)
Warning: invalid flag 0x0000 of partition table 4 will be corrected by w(rite)

Command (m for help): m
Command action
a toggle a bootable flag
b edit bsd disklabel
c toggle the dos compatibility flag
d delete a partition
l list known partition types
m print this menu
n add a new partition
o create a new empty DOS partition table
p print the partition table
q quit without saving changes
s create a new empty Sun disklabel
t change a partition's system id
u change display/entry units
v verify the partition table
w write table to disk and exit
x extra functionality (experts only)

Command (m for help): n
Command action
e extended
p primary partition (1-4)
p
Partition number (1-4): 1
First cylinder (1-5221, default 1):
Using default value 1
Last cylinder or +size or +sizeM or +sizeK (1-5221, default 5221):
Using default value 5221

Command (m for help): t
Selected partition 1
Hex code (type L to list codes): L

0 Empty 1e Hidden W95 FAT1 80 Old Minix be Solaris boot
1 FAT12 24 NEC DOS 81 Minix / old Lin bf Solaris
2 XENIX root 39 Plan 9 82 Linux swap / So c1 DRDOS/sec (FAT-
3 XENIX usr 3c PartitionMagic 83 Linux c4 DRDOS/sec (FAT-
4 FAT16 <32M 40 Venix 80286 84 OS/2 hidden C: c6 DRDOS/sec (FAT-
5 Extended 41 PPC PReP Boot 85 Linux extended c7 Syrinx
6 FAT16 42 SFS 86 NTFS volume set da Non-FS data
7 HPFS/NTFS 4d QNX4.x 87 NTFS volume set db CP/M / CTOS / .
8 AIX 4e QNX4.x 2nd part 88 Linux plaintext de Dell Utility
9 AIX bootable 4f QNX4.x 3rd part 8e Linux LVM df BootIt
a OS/2 Boot Manag 50 OnTrack DM 93 Amoeba e1 DOS access
b W95 FAT32 51 OnTrack DM6 Aux 94 Amoeba BBT e3 DOS R/O
c W95 FAT32 (LBA) 52 CP/M 9f BSD/OS e4 SpeedStor
e W95 FAT16 (LBA) 53 OnTrack DM6 Aux a0 IBM Thinkpad hi eb BeOS fs
f W95 Ext'd (LBA) 54 OnTrackDM6 a5 FreeBSD ee EFI GPT
10 OPUS 55 EZ-Drive a6 OpenBSD ef EFI (FAT-12/16/
11 Hidden FAT12 56 Golden Bow a7 NeXTSTEP f0 Linux/PA-RISC b
12 Compaq diagnost 5c Priam Edisk a8 Darwin UFS f1 SpeedStor
14 Hidden FAT16 `/mnt/disk/home'
`/home/jblanco' -> `/mnt/disk/home/jblanco'
`/home/jblanco/.bashrc' -> `/mnt/disk/home/jblanco/.bashrc'
`/home/jblanco/.bash_profile' -> `/mnt/disk/home/jblanco/.bash_profile'
`/home/jblanco/.bash_logout' -> `/mnt/disk/home/jblanco/.bash_logout'
`/home/jblanco/.sudo_as_admin_successful' -> `/mnt/disk/home/jblanco/.sudo_as_admin_successful'
`/home/jblanco/.bash_history' -> `/mnt/disk/home/jblanco/.bash_history'
`/home/jblanco/downloads' -> `/mnt/disk/home/jblanco/downloads'
jblanco@us01:~$ cd /mnt/disk
jblanco@us01:~$ cd /mnt/disk
jblanco@us01:/mnt/disk$ ls
home lost+found

GNU nano 1.3.10 File: /etc/fstab
# /etc/fstab: static file system information.
#
#
proc /proc proc defaults 0 0
/dev/sda1 / ext3 defaults,errors=remount-ro 0 1
/dev/sda5 none swap sw 0 0
/dev/sdb1 /nas ext3 defaults 0 0
/dev/hdc /media/cdrom0 udf,iso9660 user,noauto 0 0
/dev/fd0 /media/floppy0 auto rw,user,noauto 0 0

jblanco@us01:/nas$ sudo chown -R jblanco:jblanco /nas
jblanco@us01:/nas$ sudo chmod -R 755 /nas

Repository: Change from CD-ROM to Internet

There's an easy way to fix this problem. Run the following command to open the sources.list file. You can use a different editor if you feel like.

sudo vi /etc/apt/sources.list

The first thing you'll want to do is comment out this line, by placing a # symbol before it.

deb cdrom:[Ubuntu-Server 6.10 _Ed…..

You can also use this opportunity to uncomment the universe repositories.

Now that you've updated the repository list, you will have to run this command to update the local list of available software:

sudo apt-get update

Now you can apt-get install from the internet.


Packages: See What Version of a Package Is Installed on Ubuntu


dpkg -s

See Where a Package is Installed on Ubuntu

dpkg -L
dpkg -l


Disk: List disk space usage on Ubuntu

df -Th

Ubuntu Networking: From Dynamic To Static

Hercules System/370, ESA/390, and z/Architecture Emulator on Debian

Hercules is an open source software implementation of the mainframe System/370 and ESA/390 architectures, in addition to the new 64-bit z/Architecture. Hercules runs under Linux, Windows (98, NT, 2000, and XP), Solaris, FreeBSD, and Mac OS X (10.3 and later). Back in 2009 I was able to play around with the Hercules emulator and Debian zLinux/390.

$ mkdir zlinux
$ cd zlinux
$ mkdir dasd rdr prt

$ cd rdr
$ wget http://ftp.nl.debian.org/debian/dists/etch/main/installer-s390/current/images/generic/initrd.debian
$ wget http://ftp.nl.debian.org/debian/dists/etch/main/installer-s390/current/images/generic/kernel.debian
$ wget http://ftp.nl.debian.org/debian/dists/etch/main/installer-s390/current/images/generic/parmfile.debian

$ dasdinit -lfs -linux 3390.LINUX.0120 3390-3 LIN120  # /
$ dasdinit -lfs -linux 3390.LINUX.0121 3390-3 LIN121  # /home

# hercules -f s390.cnf



zlinux/s390.cnf

CPUSERIAL 000069        # CPU serial number
CPUMODEL  9672          # CPU model number
MAINSIZE  256           # Main storage size in megabytes.
XPNDSIZE  0             # Expanded storage size in megabytes
CNSLPORT  3270          # TCP port number to which consoles connect
NUMCPU    1             # Number of CPUs
LOADPARM  0120....      # IPL parameter
OSTAILOR  LINUX         # OS tailoring
PANRATE   SLOW          # Panel refresh rate (SLOW, FAST)
ARCHMODE  ESAME         # Architecture mode ESA/390 or ESAME

#
#  Device Definitions
#
# .-----------------------Device number
# |     .-----------------Device type
# |     |       .---------File name and parameters
# |     |       |
# V     V       V
#---    ----    --------------------

# console
001F    3270

# terminal
0009    3215

# reader
000C    3505    ./rdr/kernel.debian ./rdr/parmfile.debian ./rdr/initrd.debian autopad eof

# printer
000E    1403    ./prt/print00e.txt crlf

# dasd
0120    3390    ./dasd/3390.LINUX.0120
0121    3390    ./dasd/3390.LINUX.0121

# tape
0581    3420

# network                               s390     realbox
0A00,0A01  CTCI -n /dev/net/tun -t 1500 192.168.1.189 192.168.1.188



# iptables -t nat -A POSTROUTING -o eth0 -s 10.1.1.0/24 -j MASQUERADE
# iptables -A FORWARD -s 192.168.1.0/24 -j ACCEPT
# iptables -A FORWARD -d 192.168.1.0/24 -j ACCEPT
# echo 1 > /proc/sys/net/ipv4/ip_forward
# echo 1 > /proc/sys/net/ipv4/conf/all/proxy_arp



HMC:

ipl c  Note: This is device 000C - 3505

Primary Network Interface is ctc: Channel to Channel (CTC) or ESCON connection
Enter .1


Define the end-points for this virtual network interface

Select ctc device
Enter .1


Select protocol -s/390
Enter .1


Now, enter the IP addresses for the end-points (must match the IP addresses in the .cnf file).
Enter s390 box IP:

.10.1.1.2

Enter host box IP:

.10.1.1.1



Enter DNS server IP - choose the same your non-virtual system uses (see /etc/resolv.conf):

.x.x.x.x

Enter hostname:

.s390

Enter your domain name; to specify no domain name, you need to enter the empty string, but due to the way Hercules handles input, you will need to enter a dot followed by a space):

.home



Now, just sit back, and wait until your system generates a SSH key. This will take a few minutes.



Before long, the installer will ask you for a password for the remainder of the install process, just enter anything:

.foo


You'll know you are on the right track! Now, open a new terminal, and ssh into installer@10.1.1.2, if everything you did was right, ssh will ask you for a password.
Remember that you are using ssh which encrypts everything, and therefore things will be slow.
Once you enter the right password, a more familiar looking Debian installer will start up:


All done. Reboot at the end of install.

HMC: ipc 120


The system you now have is running with a 31-bit kernel. If you want a 64-bit kernel, simply run:

# aptitude install kernel-image-2.6-s390x

This will install the right image, and set up zIPL (the bootloader) to do the right thing. The original kernel image will remain installed, and you can select it in the bootloader (right after you issue ipl on the Hercules console). Enjoy!

Linux ISOS

Not really related to privacy or security, but I sometimes forget this command, and need to look it up. This way, I'll know exactly where to find it:

To make an ISO from your CD/DVD, place the media in your drive but do not mount it. If it automounts, unmount it.

dd if=/dev/dvd of=dvd.iso # for dvd
dd if=/dev/cdrom of=cd.iso # for cdrom
dd if=/dev/scd0 of=cd.iso # if cdrom is scsi


To make an ISO from files on your hard drive, create a directory which holds the files you want. Then use the mkisofs command.

mkisofs -o /tmp/cd.iso /tmp/directory/

This results in a file called cd.iso in folder /tmp which contains all the files and directories in /tmp/directory/.

Thursday, September 20, 2012

AIX Encryption Notes

Benefits:

  • The encryption is done by the OS, so no application reconfiguration is required

  • The encryption is a feature included with AIX, so it is available at no cost

  • The encryption is performed at the file system level, so encryption for main application could be phased in gradually as they have 8 distinct file systems


Concerns:

  • With the encryption being done on the server side, there are CPU cycles expended to perform the encryption

  • The estimated CPU cost is 2-5%, which is within our current idle usage availability (i.e. no additional CPUs are required in order to implement)

  • There would be downtime required to perform the encryption, but that might be able to be done multi-threaded for main application.


Unknown:

  • The AIX mechanism for encryption is geared to granting access by users and groups; in this case we would be doing group level encryption

  • The encryption key store is unique per user and has a second layer of encryption using the user’s password, so as the user changes their password, their encryption key store is re-encrypted. After the initial encryption key store has been created, an additional step is required once a user’s password changes

  • Do not yet know if processes spawned by a parent process inherit access to the encryption key store; in all likelihood they would – but if not, that would be a showstopper; however initial testing would be able to verify this almost immediately

Thursday, June 07, 2012

Paging Space Utilization

We received this alert this morning.  At the time, the paging space utilization exceeded 80%:

[yo@pd2]/>lsps -a
Page Space      Physical Volume   Volume Group Size %Used Active  Auto  Type Chksum
paging05        hdisk0            rootvg        1024MB     0    no    no    lv     0
paging04        hdisk0            rootvg        1024MB     0    no    no    lv     0
paging03        hdisk0            rootvg        1024MB     0    no    no    lv     0
paging02        hdisk0            rootvg        1024MB    80   yes   yes    lv     0
paging01        hdisk0            rootvg        1024MB    79   yes   yes    lv     0
paging00        hdisk0            rootvg        1024MB    80   yes   yes    lv     0
hd6             hdisk0            rootvg        1024MB    80   yes   yes    lv     0


There is normally 4GB of paging space active, plus there is another 3GB standing by.  Occasionally AIX paging space will become filled with stale segments, especially when java is in use on the system, which in this case is mostly from WebSphere.  By deactivating and reactivating the active paging spaces, those get cleared out.  Note: it can take up to about 10 minutes per space to deactivate.

For the procedure:
1.    Activate at least one of the spare spaces
a.    swapon /dev/paging05
2.    Step through each of the other active spaces, and issue a swapoff and swapon
a.    swapoff /dev/hd6
b.    swapon  /dev/hd6
c.    swapoff /dev/paging00
d.    swapon  /dev/paging00
e.    swapoff /dev/paging01
f.    swapon  /dev/paging01
g.    swapoff /dev/paging02
h.    swapon  /dev/paging02
3.    Release the standby space(s) activated earlier
a.    swapoff /dev/paging05

Following the procedure, we went from 80% to 5%:

[yo@pd2]/home/yo>lsps -a
Page Space      Physical Volume   Volume Group Size %Used Active  Auto  Type Chksum
paging05        hdisk0            rootvg        1024MB     0    no    no    lv     0
paging04        hdisk0            rootvg        1024MB     0    no    no    lv     0
paging03        hdisk0            rootvg        1024MB     0    no    no    lv     0
paging02        hdisk0            rootvg        1024MB     2   yes   yes    lv     0
paging01        hdisk0            rootvg        1024MB     4   yes   yes    lv     0
paging00        hdisk0            rootvg        1024MB     6   yes   yes    lv     0
hd6             hdisk0            rootvg        1024MB     9   yes   yes    lv     0


Alert:

Subject: ALERT:Warning PctTotalPgSpFree at 19.926 on pd2

Host % Total Paging Space Free PctTotalPgSpFree for pd2 triggered PctTotalPgSpFree < 20 at 19.926

Alert detail:
ERRM_DATA_TYPE=CT_FLOAT64
ERRM_RSRC_CLASS_NAME=Host
ERRM_ATTR_NAME=% Total Paging Space Free
ERRM_COND_SEVERITYID=0
ERRM_NODE_NAMELIST={pd2}
ERRM_COND_NAME=PG_Warning
ERRM_ATTR_PNAME=PctTotalPgSpFree
ERRM_COND_HANDLE=0x6004 0xffff 0xd20fd739 0x873cfd56 0x126a2686 0xfbdb1899
ERRM_RSRC_HANDLE=0x6008 0xffff 0xd20fd739 0x873cfd56 0x122601f8 0x619e2352 ERRM_TYPE=Event
ERRM_ER_HANDLE=0x6006 0xffff 0xd20fd739 0x873cfd56 0x122601f1 0x70eec021
ERRM_RSRC_NAME=pd2
ERRM_RSRC_CLASS_PNAME=IBM.Host
ERRM_ATTR_NUM=1
ERRM_TYPEID=0
ERRM_TIME=1339082122,353161
ERRM_RSRC_TYPE=0
ERRM_COND_SEVERITY=Informational
ERRM_EXPR=PctTotalPgSpFree < 20
ERRM_VALUE=19.926
ERRM_COND_BATCH=0
ERRM_NODE_NAME=pd2
ERRM_ER_NAME=Warning notifications

Wednesday, May 16, 2012

.Trashes, .fseventsd, and .Spotlight-V100

Merely plugging a removable drive into a mac (when it has write access) makes OS/X think it can take the liberty to write a lot of hidden garbage onto that disk. If you want to stop this from happening, you have to put some special files on that disk before you plug it in.

To stop OS/X from doing Spotlight indexing, you need a file called .metadata_never_index in the root directory of the removable drive.

To stop OS/X from making a .Trashes directory, you need to make your own file that *isn’t* a directory and call it .Trashes

To keep it from doing logging of filesystem events on the drive, you need to make a directory called .fseventsd and inside that folder put a single file named no_log

The contents of these files don’t matter, so you can make them empty files using touch. Even better, you could make it a text file with a link to this post, so that you (or someone else) wandering across the files will know what they’re for.

Apple’s choice to do this is incredibly self-serving and shameful. At bare minimum, hidden files and features like these should be off by default for any non-mac-only filesystem formats. They should only be enabled when the user has been made aware of them.

Tuesday, August 30, 2011

Some Helpful Tivoli Storage Manager Scripts

Here are some helpful Tivoli Storage Manager scripts:

def scr RECLAIM "select stgpool_name,reclaim from stgpools where DEVCLASS!='DISK'" desc='Reclaim Thresholds'


def scr SCRATCH "select count(*) as Scratch_count, library_name from libvolumes where status='Scratch' group by library_name" desc='Find Scratch Tape Count"

def scr DBBKSTAT "SELECT date_time, type, backup_series, volume_seq, devclass, volume_name FROM volhistory WHERE ( type='BACKUPFULL' OR type='BACKUPINCR' OR type='DBSNAPSHOT' ) AND date_time>=current_timestamp-48 hours" desc='Show Database Backup Statistics'

def scr DBCLEAN "del volh todate=-7 type=dbb" desc='Purge Database Backups older than 7 days'

def scr BACKUPDBTAPE "ba db dev=3592 t=f" desc='Database Backup to Tape'

def scr BACKUPDBDISK "ba db dev=dbbkc t=f" desc='Database Backup to Disk'

def scr NAS336 "select entity, date(start_time) as StartDate, time(start_time) as StartTime, date(end_time) as EndDate, time(end_time) as EndTime, activity, schedule_name, cast(bytes/1024/1024/1024 as decimal(18,2)) as GB, successful from summary where entity='EMCNS480' and end_time >= current_timestamp-336 hours" desc='Last 336 hours of NAS operations'

Tuesday, May 03, 2011

Information necessary to troubleshoot NAS NDMP issues

In order to minimize the overall Service Request resolution time we strongly recommend that the following information is provided and logged to the Service Request. The information will enable our Support Engineers to deal with your request in a more effective and timely manner. Please cut / paste questions plus responses into the Service Request.

  • Provide a detailed problem description that includes symptoms, error codes, error messages, and/or screen captures:


    • At what stage did the problem occur i.e. during a full backup, incremental backup, non-DAR restore, or DAR restore?

    • Was backup on a checkpoint file system or Production File System (PFS)? If it was on PFS, were files being accessed while the backup was in progress (e.g. file deletion, file manipulation or file creation etc.)?

    • Did backups or restores work properly in the past? If yes, has there been any changes whatsoever in the backup environment, tape drive, media, zoning, new param, new patches, etc..?Was the backup or restore canceled or aborted at any time?

    • How and when did the backup/restore stop?

    • When the job(s) fail what is the error message?

    • Is the backup/restore spanning to multiple tapes? If yes, are more tapes available when the job spans to the next tape?


  • NDMP client questions:


    • What is the version and patch level of the Vendor NDMP client software?

    • Is the NDMP client running on a UNIX or NT platform and what is the OS version and patch(es) or Service Pack level? What local language is the NDMP client is running?

    • Any special environmental backup settings? If yes, provide details.

    • Does the NDMP client support Direct Access Restore (DAR)? If yes, is it be enabled?

    • Is Internationalization enabled on the NDMP client?

    • Are all required processes/services running? Please verify the license and NDMP application is installed properly.

    • Please provide the NDMP software verbose log just for the problem areas if possible. For example, Networker daemons log, etc.

    • Is there an open support ticket with the Backup NDMP client vendor? If yes, please provide the ticket number and contact information. The NDMP client vendor can also assist with troubleshooting.


  • Connectivity questions:


    • Provide detailed topology information from the Data Mover to tape drive including any devices in between.

    • Please provide the device vendor and firmware version. For example, if this is SAN environment, provide switch, FC-SCSI bridge information, tape drive firmware version, etc. If it is SCSI connection, please provide the SCSI cable length.

    • Is tape library in a DDS (Dynamic Drive Sharing) environment?
      Provide topology information between Data Mover and the NDMP client vendor software host connection.

    • Please provide the IP and hostname of the NDMP Host running the NDMP software.


Thursday, April 28, 2011

CachéDB/EPIC HP Business Copy Script

Finally figured out a way to execute pairresync, pairsplit, EPIC's InstFreeze and InstThaw from the proxy backup server.

HORCMINST=1
HORCC_MRCF=1
#
logFile=/tmp/`basename $0`_`date '+%a'`.log
#
# Issue the resync
date '+%X_About to resync" >$logFile
pairresync -g dbgroup -IBC1 >>$logFile
#
# Wait for the resync to complete
#
date '+%X_Waiting for resync" >>$logFile
while true ; do
if [[ `pairdisplay -g dbgroup -IBC1 | egrep -c 'COPY|PSUS|SSUS'` -eq 0 ]] ; then
date '+%X_Resync Complete" >>$logFile
break
fi
sleep 20
done
#
sleep 60
#
# Freeze the Epic system
date '+%X_About to Freeze" >>$logFile
ssh epicserver "/epic/prd/bin/instfreeze;sync;sync;sync" >>$logFile
sleep 15
date '+%X_About to Split" >>$logFile
pairsplit -g dbgroup -IBC1 >>$logFile
sleep 15
date '+%X_About to Thaw" >>$logFile
ssh epicserver /epic/prd/bin/instthaw >>$logFile
sleep 15
date '+%X_All Done" >>$logFile

Tuesday, April 05, 2011

Backup Logs - Last 14 Days Of Any Given Server

Auditors want these backup log reports all the time. Here a short script to display the last 14 days of backup logs for any given server.

select entity, date(start_time) as StartDate, time(start_time) as StartTime, date(end_time) as EndDate, time(end_time) as EndTime, activity, cast(bytes/1024/1024/1024 as decimal(18,2)) as GB, successful from summary where entity like '%$1%' and end_time >= current_timestamp-336 hours order by entity asc

Save this code as a TSM script and run it using the following syntax:

run script-name server-name

AIX Last Login Information

for user in ` lsuser -a time_last_login ALL| grep -v =` ; do
echo “$user NEVER”
done


for user in `lsuser -a time_last_login ALL| grep =|sed 's/ time_last_login=/:/'` ; do
last=`echo $user|cut -d: -f2`
id=`echo $user | cut -d: -f1`
echo "$id `perl -le \"print scalar localtime $last \" `"
done

Wednesday, March 30, 2011

What is a JAR file?

The JAR file format is based on the popular ZIP file format, and is used for aggregating many files into one. Unlike ZIP files, JAR files are used not only for archiving and distribution, but also for deployment and encapsulation of libraries, components, and plug-ins, and are consumed directly by tools such as compilers and JVMs. Special files contained in the JAR, such as manifests and deployment descriptors, instruct tools how a particular JAR is to be treated.

A JAR file might be used:

  • For distributing and using class libraries

  • As building blocks for applications and extensions

  • As deployment units for components, applets, or plug-ins

  • For packaging auxiliary resources associated with components


The JAR file format provides many benefits and features, many of which are not provided with a traditional archive format such as ZIP or TAR. These include:

  • Security. You can digitally sign the contents of a JAR file. Tools that recognize your signature can then optionally grant your software security privileges it wouldn't otherwise have, and detect if the code has been tampered with.

  • Decreased download time. If an applet is bundled in a JAR file, the applet's class files and associated resources can be downloaded by a browser in a single HTTP transaction, instead of opening a new connection for each file.

  • Compression. The JAR format allows you to compress your files for efficient storage.

  • Transparent platform extension. The Java Extensions Framework provides a means by which you can add functionality to the Java core platform, which uses the JAR file for packaging of extensions. (Java 3D and JavaMail are examples of extensions developed by Sun.)

  • Package sealing. Packages stored in JAR files can be optionally sealed to enforce version consistency and security. Sealing a package means that all classes defined in that package must be found in the same JAR file.

  • Package versioning. A JAR file can hold data about the files it contains, such as vendor and version information.

  • Portability. The mechanism for handling JAR files is a standard part of the Java platform's core API.


Compressed and uncompressed JARs

The jar tool (see The jar tool for details) compresses files by default. Uncompressed JAR files can generally be loaded more quickly than compressed JAR files, because the need to decompress the files during loading is eliminated, but download time over a network may be longer for uncompressed files.

The jar tool

To perform basic tasks with JAR files, you use the Java Archive Tool (jar tool) provided as part of the Java Development Kit. You invoke the jar tool with the jar command. Table 1 shows some common applications:

Common usages of the jar tool

Creating a JAR file from individual files

jar -cvMf jar-file input-file...

Creating a JAR file from a directory

jar -cvMf jar-file dir-name

Creating an uncompressed JAR file

jar -cvMf0 jar-file input-file...

Updating a JAR file

jar -uf jar-file input-file

Viewing the contents of a JAR file

jar -tf jar-file

Extracting the contents of a JAR file

jar -xf jar-file

Extracting specific files from a JAR file

jar -xf jar-file archived-file


Here is the usage for AIX 6.1:

Usage: jar {ctxu}[vfm0Mi] [jar-file] [manifest-file] [-C dir] files ...
Options:
-c create new archive
-t list table of contents for archive
-x extract named (or all) files from archive
-u update existing archive
-v generate verbose output on standard output
-f specify archive file name
-m include manifest information from specified manifest file
-0 store only; use no ZIP compression
-M do not create a manifest file for the entries
-i generate index information for the specified jar files
-C change to the specified directory and include the following file
If any file is a directory then it is processed recursively.
The manifest file name and the archive file name needs to be specified
in the same order the 'm' and 'f' flags are specified.

Example 1: to archive two class files into an archive called classes.jar:
jar cvf classes.jar Foo.class Bar.class
Example 2: use an existing manifest file 'mymanifest' and archive all the
files in the foo/ directory into 'classes.jar':
jar cvfm classes.jar mymanifest -C foo/ .

Collecting Data for Tivoli Storage Manager Backup/Archive Client: Performance

Collecting Data for Tivoli Storage Manager Backup/Archive Client: Performance

Problem(Abstract)
Collecting troubleshooting documents aid in problem determination and save time resolving Problem Management Records (PMRs).



Resolving the problem
Collecting data early, even before opening the PMR, helps IBM® Support quickly determine if:
Symptoms match known problems (rediscovery).
There is a non-defect problem that can be identified and resolved.
There is a defect that identifies a workaround to reduce severity.
Locating root cause can speed development of a code fix.


Manually Gathering General Information

Please Review the following TSM Performance Tuning Guide:http://publib.boulder.ibm.com/infocenter/tivihelp/v1r1/topic/com.ibm.itsmm.doc/b_perf_tuning_guide.htm

1) From the operating system command prompt, run the following TSM command to gather up general information about TSM and the operating system environment (saved output is written to file dsminfo.txt):
dsmc QUERY SYSTEMINFO

2) From a TSM Admin command line client, enter the following commands:
QUERY SYSTEM > querysys.out

3) Gather the following file and information:
•dsminfo.txt
•querysys.out
•details of operating system levels ( i.e., HP-UX 11.23)
•TSM Client specific version (i.e., 5.4.0.2)


Manually Gathering Client Performance Specific Information

Collect TSM Performance Instrumentation Traces:

1) From the poor performing TSM client system, issue this command from an operating system command prompt:
dsmc INCREMENTAL -TESTFLAG=INSTRUMENT:DETAIL > backup.out
Note:
•-TESTFLAG=INSTRUMENT:DETAIL will generate a file with name "dsminstr.report.x "under the same directory where dsmerror.log locates. For Netware, the file name will be called dsminstr.rep.
•If the slow client is an API client, replace -TESTFLAG=INSTRUMENT:DETAIL with -TESTFLAG=INSTRUMENT:API
•If INCREMENTAL is not the slow action, replace dsmc INCREMENTAL with the appropriate slow TSM client command.
•For users who can't run a manual backup, add the following lines in the according option file for example: dsm.opt
TESTFLAG INSTRUMENT:DETAIL
Recycle Scheduler to active the trace

2) Collect a TSM server instrument trace when the slow client backup is running:
From a TSM Admin command line client, enter the following commands and collect output and trace:
a) Start Instrument tracing:
INSTrument BEGIN (leave trace running for 20 minutes)
b) Collect Process and Session output:
QUERY SESS > sess.out
QUERY PROC > proc.out
QUERY DB > db.out

c) End Instrument tracing:
INSTrument END FILE=inst.out
3) Gather the following files :
•backup.out
•dsminstr.report.x (dsminstr.rep on Netware)
•sess.out
•proc.out
•db.out
•inst.out

Tuesday, March 29, 2011

Counting TSM Sessions Every 15 Minutes and Logging To File With Time Stamp

#!/bin/ksh
#
SessCnt=`dsmadmc -id=admin -password=password -dataonly=yes q sess| grep -c Node`
#
date "+%y%m%d:%X_TSM Session Count=$SessCnt" >>/tmp/tsm_sessions
if [[ $SessCnt -gt 100 ]] ; then
echo TSM Session Count is $SessCnt | mail -s TSM_Sessions_High aix@company.com
fi
#
if [[ $SessCnt -gt 185 ]] ; then
echo TSM Session Count is $SessCnt | mail -s TSM_Sessions_at_MAX joe1@company.com joe2@company.com joe3@company.com
echo TSM Session Count is $SessCnt | mail -s TSM_Sessions_at_MAX joe1@remote.com
fi


Then create a crontab entry

0,15,30,45 * * * * /usr/local/scripts/tsm_session_ck

Monday, March 21, 2011

Delete All Logical Volumes in a Volume Group

for  lv  in  `lsvg  -l  <Volume_Group_Name>  |  grep /  |  awk '{print $1}'`  ;  do
echo  rmlv  -f  $lv
done


Remove the echo to execute.

Wednesday, December 15, 2010

10 PowerShell commands every Windows admin should know

Over the last few years, Microsoft has been trying to make PowerShell the management tool of choice. Almost all the newer Microsoft server products require PowerShell, and there are lots of management tasks that can’t be accomplished without delving into the command line. As a Windows administrator, you need to be familiar with the basics of using PowerShell. Here are 10 commands to get you started.

Note: This article is also available as a PDF download.

1: Get-Help


The first PowerShell cmdlet every administrator should learn is Get-Help. You can use this command to get help with any other command. For example, if you want to know how the Get-Process command works, you can type:
Get-Help -Name Get-Process

and Windows will display the full command syntax.

You can also use Get-Help with individual nouns and verbs. For example, to find out all the commands you can use with the Get verb, type:
Get-Help -Name Get-*

2: Set-ExecutionPolicy


Although you can create and execute PowerShell scripts, Microsoft has disabled scripting by default in an effort to prevent malicious code from executing in a PowerShell environment. You can use the Set-ExecutionPolicy command to control the level of security surrounding PowerShell scripts. Four levels of security are available to you:

  • Restricted — Restricted is the default execution policy and locks PowerShell down so that commands can be entered only interactively. PowerShell scripts are not allowed to run.

  • All Signed — If the execution policy is set to All Signed then scripts will be allowed to run, but only if they are signed by a trusted publisher.

  • Remote Signed — If the execution policy is set to Remote Signed, any PowerShell scripts that have been locally created will be allowed to run. Scripts created remotely are allowed to run only if they are signed by a trusted publisher.

  • Unrestricted — As the name implies, Unrestricted removes all restrictions from the execution policy.


You can set an execution policy by entering the Set-ExecutionPolicy command followed by the name of the policy. For example, if you wanted to allow scripts to run in an unrestricted manner you could type:
Set-ExecutionPolicy Unrestricted

3: Get-ExecutionPolicy


If you’re working on an unfamiliar server, you’ll need to know what execution policy is in use before you attempt to run a script. You can find out by using the Get-ExecutionPolicy command.

4: Get-Service


The Get-Service command provides a list of all of the services that are installed on the system. If you are interested in a specific service you can append the -Name switch and the name of the service (wildcards are permitted) When you do, Windows will show you the service’s state.

5: ConvertTo-HTML


PowerShell can provide a wealth of information about the system, but sometimes you need to do more than just view the information onscreen. Sometimes, it’s helpful to create a report you can send to someone. One way of accomplishing this is by using the ConvertTo-HTML command.

To use this command, simply pipe the output from another command into the ConvertTo-HTML command. You will have to use the -Property switch to control which output properties are included in the HTML file and you will have to provide a filename.

To see how this command might be used, think back to the previous section, where we typed Get-Service to create a list of every service that’s installed on the system. Now imagine that you want to create an HTML report that lists the name of each service along with its status (regardless of whether the service is running). To do so, you could use the following command:
Get-Service | ConvertTo-HTML -Property Name, Status > C:\services.htm

6: Export-CSV


Just as you can create an HTML report based on PowerShell data, you can also export data from PowerShell into a CSV file that you can open using Microsoft Excel. The syntax is similar to that of converting a command’s output to HTML. At a minimum, you must provide an output filename. For example, to export the list of system services to a CSV file, you could use the following command:
Get-Service | Export-CSV c:\service.csv

7: Select-Object


If you tried using the command above, you know that there were numerous properties included in the CSV file. It’s often helpful to narrow things down by including only the properties you are really interested in. This is where the Select-Object command comes into play. The Select-Object command allows you to specify specific properties for inclusion. For example, to create a CSV file containing the name of each system service and its status, you could use the following command:
Get-Service | Select-Object Name, Status | Export-CSV c:\service.csv

8: Get-EventLog


You can actually use PowerShell to parse your computer’s event logs. There are several parameters available, but you can try out the command by simply providing the -Log switch followed by the name of the log file. For example, to see the Application log, you could use the following command:
Get-EventLog -Log "Application"

Of course, you would rarely use this command in the real world. You’re more likely to use other commands to filter the output and dump it to a CSV or an HTML file.

9: Get-Process


Just as you can use the Get-Service command to display a list of all of the system services, you can use the Get-Process command to display a list of all of the processes that are currently running on the system.

10: Stop-Process


Sometimes, a process will freeze up. When this happens, you can use the Get-Process command to get the name or the process ID for the process that has stopped responding. You can then terminate the process by using the Stop-Process command. You can terminate a process based on its name or on its process ID. For example, you could terminate Notepad by using one of the following commands:
Stop-Process -Name notepad
Stop-Process -ID 2668

Keep in mind that the process ID may change from session to session.

Monday, November 29, 2010

Top 8 Traits Employers Look For

When looking for employment you should keep in mind the needs of potential employers. Of course you want the job – but what do employers want from you?

1. Loyalty – Loyal employees are the backbone of any organization. Look at your resume (or LinkedIn profile), does your work history seem like that of a loyal employee? If there are short term roles listed make it clear they were temporary jobs. Jobs not relevant to the role can be omitted. At interview do not gossip or share details of previous employers, colleagues or mutual acquaintances – indulging in tittle-tattle will make you appear disloyal.

2. Honesty – Employers need to be able to trust you. Never pad you resume, be honest about qualifications and experience. Nothing is more likely to make you appear dishonest than being caught in a half truth at interview.

3. Punctuality – Bosses need to know their workers will be at their posts on time every day as poor timekeeping can cost a company customers. Demonstrate good timekeeping by turning in your application form on time. If given an interview appointment make sure you arrive at least ten minutes in advance of you allotted time.

4. Determination – People who want to do well are much more likely to work with passion and gain results for their employers. When given the chance to question a potential employer at interview, don’t be afraid to ask about opportunities for advancement, as this will display determination and alert the interviewer to your desire to succeed.

5. Flexibility – The ever changing face of business means employers need workers who are willing to move with the times, adapting to the evolving demands of their role. Use your application form to demonstrate how you have been flexible in previous roles – if you took on extra responsibilities or undertook additional training make sure you let them know.

6. Smart Appearance – Whether you will be working in a formal or casual environment, it is important to take care with your appearance at interview. If you look scruffy employers may ignore your skills, assuming that someone who doesn’t take pride in themselves will be unlikely to take pride in their work.

7. Positive Outlook – Employers want go-getters on their team. Showing a positive outlook at interview is essential. If you have been made redundant from a previous role, don’t bemoan your fate to a potential boss, instead explain how excited you are to be exploring new opportunities.

8. Communications Skills – Communication skills are key. Ensure your application and resume are grammatically correct, ask a friend or family to look them over for you to help catch any typos or errors. At interview think before you speak, choose your words carefully and ensure they convey the message clearly. Don’t “um” and “ah” – it is better to take a few seconds to compose an answer than to say something stupid.

Demonstrating these eight key skills can help you land your dream role, and displaying these traits in your daily working life will make you stand out over other employees when the opportunity for advancement arises.

Saturday, November 20, 2010

Filesystem Utilization Korn Shell Script

Script will run in cron @ 6,8,10,12am and 2,4pm to check filesystems that are above 90% full.
95% is the percentage used across the board unless otherwise noted.
#!/bin/ksh
#######################################################################################################
# Created: Javier Blanco
# Creation Date: 11/06/2001
# Use: By all unix servers
# Description: Script will run in cron @ 6,8,10,12am and 2,4pm to check filesystems that are
# above 90% full.
# 95% is the percentage used across the board unless otherwise noted.
#
########################################################################################################
export AD1=unixadmins@xyz.com

export PG1=555555555@pager.net
export PG2=555555555@pager.net

df -k | grep -iv -e filesystem -e /var/nim/resources -e proc -e cd -e /var/mksysbs -e /usr | awk '{ print $7" "$4}' | while read LINE
do
FILESYSTEM=`echo $LINE | cut -d"%" -f1 | awk '{ print $1 }'`
PERCENTAGE=`echo $LINE | cut -d"%" -f1 | awk '{ print $2 }'`
if [ $PERCENTAGE -ge 95 ]; then

echo "The system `hostname` has caused a filesystem alert on `date`. The ${FILESYSTEM} is ${PERCENTAGE}% full. \\n You will recieve this message every 2 hours until either the problem is corrected or the root crontab entry is commented out.[dfchk.sh] \\n \\n " | mail -s "- $FILESYSTEM is at $PERCENTAGE " $AD1 $PG1 #$PG2
fi
done