Showing posts with label Linux. Show all posts
Showing posts with label Linux. Show all posts

Thursday, July 19, 2018

Oracle Universal Installer... Checking swap space: 493 MB available, 500 MB required. Failed

After installing the patch i decided to relink all binaries. The relinking procedure is very simple:

$ cd $OH/bin
$ ./relink all 

But in my environment i got a unpleasant message:

" Starting Oracle Universal Installer...

Checking swap space: 493 MB available, 500 MB required.    Failed <<<<

Some requirement checks failed. You must fulfill these requirements before
continuing with the installation,

Exiting Oracle Universal Installer ...  "


Because I'm on the Linux, then the fastest solution is to make additional swap file.
It can be done online:

# free
# df -h
# dd if=/dev/zero of=/swapfile bs=1M count=100
# chmod 0600 /swapfile
# mkswap /swapfile
# swapon /swapfile
# df -h
# free


And relink  say:

"Starting Oracle Universal Installer...

Checking swap space: must be greater than 500 MB.   Actual 593 MB    Passed
Preparing to launch Oracle Universal Installer ..."



To revert changes (online): 
#  swapoff /swapfile
#  rm /swapfile






Tuesday, March 20, 2018

"no space left on device"

What you should to do if you got a message "no space left on device" on some file system (/u01 for example) ?


1.

Start from inode analysis:
# date
# df -i

and look the values Inodes, IUsed, IFree, IUse% in the output.

If all 100% inode are in use, then we suspect a lot of files were created recently in one of sub-directories.
To find biggest (by number of files) directory :
# find /u01 -xdev -printf '%h\n' | sort | uniq -c | sort -k 1 -n | tail -50
You'll got a 50 biggest directories.
Delete redundant files.
Try to do it without rm command, use adrci or find with -mtime 30 option.
Deleteing many files require many time, so run these commands in screen.

Adrci example:

[oracle@edbadm01 trace]$ adrci
adrci> show homes
adrci> set home diag/rdbms/irbis/irbis1
adrci> purge -age 60
       - Purge diagnostic data older than <mins> from the ADR home
Or
adrci> purge -size 10485760000 - Purge diagnostic data from the ADR home until the size of the home reaches <bytes> bytes (approx 10G).

Adrci can remove any number of files without annoing "/bin/rm: Argument list too long" message.
Also this program will work after your terminal window is accidentally closed (or network failure).
But adrci is command from Oracle Home and it can work with Oracle trace files only.

Find example:
use find and -mtime option:
# date; find /u01/app -name *aud -mtime +30 -type f -delete
3rd way: rm. It is the least recommended way because you should to use rm to delete closed files only.  Don't use rm to remove files opened by any process. Test files with fuser command before removing.
It's very easy to make a mistake using rm, so use adrci or find if possible.


2.

If the previous step found many free inodes (ensured there are no obstacles to create a file)
then we need to investigate free space in file system:
# date; df -h

We need to understand how many free space has interestiong file system.
If there is no free space ( Size = Used, Use% = 100% ), then we need to delete some files.

Start analyze biggest directories (by sum of file size ):
# date; du -m /u01/app/|sort -rn |less 

Usually i choose the 1st biggest directory and delete files in it.
Then i choose next biggest directory and so on ... until i have enough space.

Usually i use 2 command to navigate through the catalog tree :
# du -sm * | sort -rn  - to understand which sub-directory is biggest
and
# cd  some_dir - to go to this directory

It may be useful to quickly find big files in your file system:
#  find /u01 -size +50M
Don't forget about -mtime option, it is good addition to point old files.


3.

If you've found enough number of free inodes & enough quantity of free space
but you still get "no space left on device" then this means that some process
has big opened file & this file was removed by rm command & but this process still running and keep this file opened.

Hint:
If the size of big file should be set to zero, then use > command (not rm this file) :

[oracle@ed04dbadm01 log]$ cd /u01/app/oracle/diag/rdbms/killme122/km1221/trace/
[oracle@ed04dbadm01 trace]$ ls -l alert_km1221.log
-rw-r----- 1 oracle asmadmin 430376 Mar 20 08:00 alert_km1221.log
[oracle@ed04dbadm01 trace]$ > alert_km1221.log
[oracle@ed04dbadm01 trace]$ ls -l alert_km1221.log
-rw-r----- 1 oracle asmadmin 0 Mar 20 11:11 alert_km1221.log


If you use rm to delete big files then you will get unexplained loss of disk space.

To know which files are opened use lsof command (list of open files).
It shows all open files for all processes in system, PIDs and names of processes.
Its output produce many lines on the screen so i use lsof | less .

Analyse open files - is difficult case, but you can solve this problem easy and quickly if it is possible
- to kill process[es] (kill -9) which have opened files in our file system (# lsof | grep /u01/app )
  or
- to restart or shutdown process[es] which have opened files in our file system (# lsof | grep /u01/app )
  or
- to restart OS (operation system or server).

After this you will see free space in file system.

4.
If you were not helped by previous recommendations then take new disk and extend your file system ;)
It may be not a joke.
If you inspect all your files and found all files as necessary (no extra files and no free space)
then it is become a time to grow this file system.

Good luck !

Tuesday, October 7, 2014

timed out waiting for input: auto-logout

I connected to db node of Exadata and after some idle time I was disconnected from server with the message:

[root@ed02dbadm02 ~]# timed out waiting for input: auto-logout

Connection closed.

As I remember my terminal can stay connected to Exadata all day, many hours.  
Whats the matter ? It is something new ...


Internet explain this behaviour:

Auto-Logout Timeout in SSH


The ssh "timed out waiting for input: auto-logout" messages is generated by ssh upon reaching a auto-logout after an inactivity time specified by the TMOUT environment variable. If this variable is not set your session will not be auto-logged out due to inactivity. If the environment variable is set, your session will be automatically closed/logged out after the amount of seconds specified by the TMOUT variable.

To see if your auto-logout variable is set and/or see what it is set to issue the following command:
 $ echo $TMOUT

Often this value is defined in /etc/profile (globally) or your user's profile (~/.profile or ~/.bash_profile).

To alter the auto-logout amount, set the TMOUT environment variable accordingly:
* TMOUT=600   #set an auto-logout timeout for 10 minutes
* TMOUT=1200  #set an auto-logout timeout for 20 minutes
* TMOUT=   #turn off auto-logout (user session will not auto-logout due to session inactivity)

This value can be set globally (e.g. TMOUT=1200) in the /etc/profile file;

Saturday, August 2, 2014

The hidden space in Exadata, part 1

The very few people are aware about unallocated (free) space on the local disks of DB nodes of Exadata.

Connect to db node, run the vgdisplay and look the "Free PE" line. I'll show the X2-2 example:

[root@dbm1db02 ~]# vgdisplay
  --- Volume group ---
  VG Name               VGExaDb
  System ID            
  Format                lvm2
  Metadata Areas        1
  Metadata Sequence No  13
  VG Access             read/write
  VG Status             resizable
  MAX LV                0
  Cur LV                3
  Open LV               3
  Max PV                0
  Cur PV                1
  Act PV                1
  VG Size               556.80 GB
  PE Size               4.00 MB

  Total PE              142541
  Alloc PE / Size       39424 / 154.00 GB
  Free  PE / Size       103117 / 402.80 GB



As you can see we have 103117 free Physical Extent (PE Size =  4.00 MB).

Let use it:

[root@dbm1db02 ~]# lvcreate -l 103117 -n LVStore VGExaDb
  Logical volume "LVStore" created


I created the  " LVStore" - new logical volume in "VGExaDb" disk group.

You can see it in another way:

[root@dbm1db02 ~]# ls -l /dev/VGExaDb/
lrwxrwxrwx 1 root root 28 Jun 11 22:30 LVDbOra1 -> /dev/mapper/VGExaDb-LVDbOra1
lrwxrwxrwx 1 root root 29 Jun 11 22:30 LVDbSwap1 -> /dev/mapper/VGExaDb-LVDbSwap1
lrwxrwxrwx 1 root root 28 Jun 11 22:30 LVDbSys1 -> /dev/mapper/VGExaDb-LVDbSys1
lrwxrwxrwx 1 root root 27 Aug  2 16:57 LVStore -> /dev/mapper/VGExaDb-LVStore


The disk group VGExaDb now have no free space:

[root@dbm1db02 ~]# vgdisplay
  --- Volume group ---
  VG Name               VGExaDb
  System ID            
  Format                lvm2
  Metadata Areas        1
  Metadata Sequence No  14
  VG Access             read/write
  VG Status             resizable
  MAX LV                0
  Cur LV                4
  Open LV               3
  Max PV                0
  Cur PV                1
  Act PV                1
  VG Size               556.80 GB
  PE Size               4.00 MB
  Total PE              142541
  Alloc PE / Size       142541 / 556.80 GB
  Free  PE / Size       0 / 0


Let use new volume:

[root@dbm1db02 ~]# mkfs -t ext3 /dev/mapper/VGExaDb-LVStore
 

mke2fs 1.39 (29-May-2006)
Filesystem label=
OS type: Linux
Block size=4096 (log=2)
Fragment size=4096 (log=2)
52805632 inodes, 105591808 blocks
5279590 blocks (5.00%) reserved for the super user
First data block=0
Maximum filesystem blocks=4294967296
3223 block groups
32768 blocks per group, 32768 fragments per group
16384 inodes per group
Superblock backups stored on blocks:
    32768, 98304, 163840, 229376, 294912, 819200, 884736, 1605632, 2654208,
    4096000, 7962624, 11239424, 20480000, 23887872, 71663616, 78675968,
    102400000

Writing inode tables: done                           
Creating journal (32768 blocks): done
Writing superblocks and filesystem accounting information:
done

This filesystem will be automatically checked every 29 mounts or
180 days, whichever comes first.  Use tune2fs -c or -i to override.
[root@dbm1db02 ~]#
[root@dbm1db02 ~]# mkdir /store
[root@dbm1db02 ~]# e2label /dev/mapper/VGExaDb-LVStore store
[root@dbm1db02 ~]# vi /etc/fstab
[root@dbm1db02 ~]# mount /store 
[root@dbm1db02 ~]# df -h /store
Filesystem            Size  Used Avail Use% Mounted on
/dev/mapper/VGExaDb-LVStore
                      397G  199M  377G   1% /store


As you can see the X2-2 users can have additional 377g as new file system.

On the X4-2 Exadata you will have about 1.5T of additional space:

[root@ed02dbadm01 ~]# vgdisplay
  --- Volume group ---
  VG Name               VGExaDb
  System ID            
  Format                lvm2
  Metadata Areas        2
  Metadata Sequence No  9
  VG Access             read/write
  VG Status             resizable
  MAX LV                0
  Cur LV                5
  Open LV               4
  Max PV                0
  Cur PV                2
  Act PV                2
  VG Size               1.63 TB
  PE Size               4.00 MB
  Total PE              428308
  Alloc PE / Size       428308 / 1.63 TB
  Free  PE / Size       0 / 0


[root@ed02dbadm01 ~]# df -h
Filesystem            Size  Used Avail Use% Mounted on
/dev/mapper/VGExaDb-LVStore
                      1.5T 1017G  375G  74% /store



Or you can go other way and extend / or /u01 using this unallocated space.


Tuesday, March 11, 2014

ASYNC IO is not working in 11.2.0.4.3 (January 2014 quarterly full stack patch)

After applying to the Exadata the January 2014 quarterly full stack patch and 11.2.0.4.3 RDBMS we run the IO Calibration dbms_resource_manger.calibrate_io(). It generated the messages in alert log trace files:

calibrate_io (kcfcagmdf): WARNING: ASync I/O not possible for datafile with file number (fno) [1]
calibrate_io (kcfcagmdf): WARNING: ASync I/O not possible for datafile with file number (fno) [2]
calibrate_io (kcfcagmdf): WARNING: ASync I/O not possible for datafile with file number (fno) [3]
calibrate_io (kcfcagmdf): WARNING: ASync I/O not possible for datafile with file number (fno) [4]


We checked the DB parameters:

SQL> show parameter filesystemio

NAME                 TYPE   VALUE
-------------------- ------ ------
filesystemio_options string setall

SQL> show parameter disk_asynch_io

NAME           TYPE    VALUE
-------------- ------- -------
disk_asynch_io boolean TRUE


We checked the  ASYNC IO inside the DB:

SQL> COL NAME FORMAT A50
SELECT NAME,ASYNCH_IO FROM V$DATAFILE F,V$IOSTAT_FILE I
WHERE F.FILE#=I.FILE_NO
AND FILETYPE_NAME='Data File';

NAME                                          ASYNCH_IO
--------------------------------------------- ---------
+DATAC1/migdb/datafile/system.256.839351447   ASYNC_OFF
+DATAC1/migdb/datafile/sysaux.257.839351449   ASYNC_OFF
+DATAC1/migdb/datafile/undotbs1.258.839351449 ASYNC_OFF
+DATAC1/migdb/datafile/users.259.839351449    ASYNC_OFF


Then we checked the libaio libraries in Linux (they are well):

[root@dm01dbadm01 ~]# rpm -qa|grep libaio
libaio-0.3.106-5
libaio-devel-0.3.106-5
libaio-devel-0.3.106-5
libaio-0.3.106-5
[root@dm01dbadm01 ~]# find / -name libaio*
/usr/lib64/libaio.so.1
/usr/lib64/libaio.so.1.0.0
/usr/lib64/libaio.so.1.0.1
/usr/lib64/libaio.a
/usr/lib64/libaio.so
/usr/lib/libaio.so.1.0.0
/usr/lib/libaio.so.1
/usr/lib/libaio.so.1.0.1
/usr/lib/libaio.a
/usr/lib/libaio.so
/usr/share/doc/libaio-0.3.106
/usr/include/libaio.h
/u01/app/oracle/product/11.2.0.4/dbhome_1/lib/stubs/libaio.so.1
/u01/app/oracle/product/11.2.0.4/dbhome_1/lib/stubs/libaio.so
/u01/app/oracle/product/11.2.0.4/dbhome_1/lib/stubs/libaio-2.3.4-stub.so
/u01/app/11.2.0.4/grid/lib/stubs/libaio.so.1
/u01/app/11.2.0.4/grid/lib/stubs/libaio.so
/u01/app/11.2.0.4/grid/lib/stubs/libaio-2.3.4-stub.so


And the async syscall statistics:

# cat /proc/slabinfo|grep kio
kioctx 86 110 384 10 1 : tunables 54 27 8 : slabdata 11 11 0
kiocb 0 0 256 15 1 : tunables 120 60 8 : slabdata 0 0 0
# more /proc/sys/fs/aio-max-nr
3145728
# more /proc/sys/fs/aio-nr
11008


So, we understand - we loosed the ASYNC IO in the Exadata.

Solution:  Patch 16618055 :

The fix Patch 16618055 : PHSB: EXADATA DOESN'T SUPPORT ASYNC I/O is now available for download on top of 11.2.0.4.3

This patch is online installable, we installed it.
In the alert log we see:
Tue Mar 11 13:15:02 2014
Patch bug16618055.pch Installed - Update #1
Patch bug16618055.pch Enabled - Update #2
Tue Mar 11 13:15:05 2014
Online patch bug16618055.pch has been installed
Online patch bug16618055.pch has been enabled
 

But the DB showed ASYNC_OFF
NAME                                          ASYNCH_IO
--------------------------------------------- ---------
+DATAC1/migdb/datafile/system.256.839351447   ASYNC_OFF
+DATAC1/migdb/datafile/sysaux.257.839351449   ASYNC_OFF
+DATAC1/migdb/datafile/undotbs1.258.839351449 ASYNC_OFF
+DATAC1/migdb/datafile/users.259.839351449    ASYNC_OFF


After shutdown the ASYNC become ON !
 

Thursday, February 13, 2014

Patchmgr hungs in the process of patching IB switches

As you know, the January 2014 Quarterly Full Stack Patch (QFSP) for Exadata contain new IB switches version - 2.1.3-4.
In many cases (my experience - 4 switches of 6) the command hungs:
./patchmgr -ibswitches -upgrade

It nothing do and wait something. We waited 1 hour and press Ctrl+C.
 
The SOLUTION is:

When you connect to Unix server the ssh daemon running on this server get your IP and try to resolve your IP into the name.
Such behavior is controlled by parameter UseDNS in /etc/ssh/sshd_config file.
By default the line look like:
#UseDNS=yes

Remove the # and change it to "no" :
UseDNS=no

And run patchmgr again !

Friday, January 31, 2014

System performance utilities show wrong CPU load on Intel processosrs when HyperThreading is enables

The interesting information is published in the

 
pages 28-29:



"Compute node CPU utilization can be measured through many different tools – top, AWR, iostat, vmstat, etc. and they all give the same number, and % CPU utilization typically averaged over a set period of time. Choose whichever tool is most convenient, but allow for Intel CPU hyper-threading.
The Intel CPUs used in all Exadata models run with two threads per CPU core. This helps to boost overall performance, but the second thread is not as powerful as the first. The operating system assumes that all threads are equal thus overstates the available CPU capacity by the operating system. We need to allow for this. Here is an approximate rule of thumb that can be used to estimate actual CPU utilization, but note that this can vary with different workloads:
∙ For CPU utilization less than 50%, multiply by 1.7.
∙ For CPU utilization over 50%, assume 85% plus (util-50%)* 0.3.

Here is a table that summarizes the effect:
Measured Utilization


Actual Utilization
10%
17%
20%
34%
30%
51%
40%
68%
50%
85%
60%
88%
70%
91%
80%
94%
90%
97%
100%
100%
 "

This information is applicable to all x86 abd x86-64 servers with HyperThreading enabled.

Enabling HT causes system statistics tools - vmstat, sar - show incorrect CPU load,
and Oracle performance tools - AWR, ASH - also show and store incorrect CPU load statistics.

Be careful !



Friday, September 20, 2013

Command: ChgPassWordString grid welcome1 produced null output


Making new installation of 1/2 Exadata i've got a message:

[root@dm01dbadm01 linux]# ./install.sh -cf ./WorkDir/MegaYEKTTest.xml -s 3

20 Sep 13 14:19:45 [INFO ] Executing Create Users
20 Sep 13 14:19:45 [INFO ] Creating users...
20 Sep 13 14:19:45 [INFO ] Creating users in cluster cluster-clu1 ................................
20 Sep 13 14:20:23 [INFO ] Following errors were found while checking command output:
20 Sep 13 14:20:23 [INFO ] ERROR:
20 Sep 13 14:20:23 [INFO ] Command: ChgPassWordString grid welcome1 produced null output but executed successfully on dm01dbadm03
20 Sep 13 14:20:23 [INFO ] zipping log and WorkDir directories . . 
20 Sep 13 14:20:23 [INFO ] Please send /opt/oracle.SupportTools/linux/WorkDir/Diag-130920_142023.zip to Oracle if you require assistance...
20 Sep 13 14:20:23 [INFO ] OcmdException from node dm01dbadm01.mega.com return code = 2 output string: Error running command ChgPassWordString grid welcome1 on node dm01dbadm03
20 Sep 13 14:20:23 [INFO ] OcmdException from node dm01dbadm01.mega.com return code = 2 output string: Error running Create Users error message Error running
oracle.onecommand.deploy.users.DeployUserUtils method createAllUsers






I logined to node 3 and 4 and tried to change password manually, and confirmed the error:

[root@dm01dbadm04 ~]# passwd
Changing password for user root.
... some silent seconds and .
passwd: Authentication token manipulation error

But in node 1 and node 2 passwd worked well.

I did
# strace passwd
and noticed all files passwd is open
open("/etc/pam.d/system-auth", O_READONLY) =...
They are
/etc/passwd, /etc/shadow, /etc/pam.d/passwd, /etc/pam.d/system-auth, /etc/pam.d/other

 Then I compared  these files at good and bad nodes.
Actually the difference was in some commented lines in /etc/pam.d/system-auth in bad node.  

 SOLUTION


I copied files from good node to bad nodes:
/etc/passwd, /etc/shadow, /etc/pam.d/passwd, /etc/pam.d/system-auth, /etc/pam.d/other


And the problem was gone :) !

Monday, May 13, 2013

error: Zip file too big (greater than 4294959102 bytes)

The customer send us the dump for Exadata POC and this dump is unavailable to unzip :

[root@ed01db01 ]# unzip dmp.zip
error:  Zip file too big (greater than 4294959102 bytes)
Archive:  dmp.zip
warning [dmp.zip]:  10149171547 extra bytes at beginning or within zipfile
  (attempting to process anyway)
error [dmp.zip]:  start of central directory not found;
  zipfile corrupt.
  (please check that you have transferred or created the zipfile in the
  appropriate BINARY mode and that you have compiled UnZip properly)

Our unzip version is:

[root@ed01db01 ]# unzip
UnZip 5.52 of 28 February 2005, by Info-ZIP.  Maintained by C. Spieler.  Send
bug reports using
http://www.info-zip.org/zip-bug.html; see README for details.

The solution 1:

[root@ed01db01 ttelek]# funzip dmp.zip > dmp.dmp
[root@ed01db01 ttelek]# ls -l

-rw-r--r-- 1 root   root 65195868160 May 13 11:24 dmp.dmp
-rw-r----- 1 oracle dba  14444138961 May 12 00:03 dmp.zip

[root@ed01db01 ttelek]# funzip
fUnZip (filter UnZip), version 3.94 of 17 February 2002
usage: ... | funzip [-password] | ...
       ... | funzip [-password] > outfile
       funzip [-password] infile.zip > outfile
       funzip [-password] infile.gz > outfile
Extracts to stdout the gzip file or first zip entry of stdin or the given file.

Our conclusion: the Exadata has the funzip by default, we don't install it intentionally.

The solution 2:

[root@ed01db01 ]# zcat dmp.zip > dmp.dmp

Exadata has zcat by default.

 

Tuesday, March 19, 2013

oracleasm createdisk [FAILED]

The today error was:

# /etc/init.d/oracleasm createdisk VOL1 /dev/sdb1
Marking disk "VOL1" as an ASM disk:               [FAILED]


What is the reason for  [FAILED] ?
The disks are:
# ls -l /dev/sdb*
brw-rw----. 1 root disk 8, 16 Mar 19 11:47 /dev/sdb
brw-rw----. 1 root disk 8, 17 Mar 19 11:58 /dev/sdb1



The solution is DISABLE SELINUX !

The dot in the brw-rw----. <- is the symptom of selinux.

After SELINUX=disabled in /etc/selinux/config +reboot the server we have

# ls -l /dev/sdb*
brw-rw---- 1 root disk 8, 16 Mar 19 13:52 /dev/sdb
brw-rw---- 1 root disk 8, 17 Mar 19 14:01 /dev/sdb1


and oracleasm configuration works:

# /etc/init.d/oracleasm createdisk VOL1 /dev/sdb1
Marking disk "VOL1" as an ASM disk:                [  OK  ]


 

Thursday, September 27, 2012

You (oracle) are not allowed to use this program (crontab)


Сегодня выяснилось, что в Экзадате не удается создать задание для cron:

[root@ed01db01 etc]# su - oracle
[oracle@ed01db01 ~]$ crontab -l
You (oracle) are not allowed to use this program (crontab)
See crontab(1) for more information

Причем, эта беда наблюдается не только под пользователем oracle, но и под другими:

[root@ed01db01 pam.d]# useradd yu
[root@ed01db01 pam.d]# su - yu
[yu@ed01db01 ~]$ crontab -l
You (yu) are not allowed to use this program (crontab)

Металинк дает свои дельные советы, которые тоже не работают :
 
Non Root Users Are Unable To Create Crontab Entries [ID 1268765.1]
Cron Scripts Might Not Be Executed After A Fresh Install On Exadata [ID 1323999.1]
 
 
Все оказалось примитивно просто: для решения проблемы достаточно добавить oracle в /etc/cron.allow  !
 
Вспоминаем основы:
 
You can execute crontab if your name appears in the file /etc/cron.allow.
If that file does not exist, you can use crontab if your name does not appear in the file /etc/cron.deny.
If only cron.deny exists and is empty, all users can use crontab.
If neither file exists, only the root user can use crontab.
 

Monday, May 21, 2012

How do register an Exadata with Unbreakable Linux Network (ULN)?

I am doing the installation of last FSQDPE patch 11.2.3.1.0 at April 2012.

Unlike previous patches the new the 11.2.3.1.0 Exadata Storage Server patch requires that the OS Linux on DB  nodes will be updated.  In section "6.1 Updating Oracle Linux Database Servers in Oracle Exadata Database Machine X2-2" Oracle say to register an Exadata with ULN.
My system is fresh (installed in March 2012) and have not been registered in ULN and have no patches yet.

So I get the error for YUM command:

[root@ed01db01 Server]# yum --enablerepo=exadata_dbserver_11.2.3.1.0_x86_64_base repolist
Error getting repository data for exadata_dbserver_11.2.3.1.0_x86_64_base, repository not found

The YUM Repository Setup give the advice:
Register the machine on the Unbreakable Linux Network with commands:
rpm --import /usr/share/rhn/RPM-GPG-KEY
up2date-nox --register

But there are no RPM-GPG-KEY file and up2date command in the fresh Exdata!

Another step was ound in Metalink:
NOTE 1234710.1, "Enable database hosts in Exadata Database Machine for using up2date or yum and vncserver"

This note offers to run this command to install up2date:

rpm -Uhv --nodeps curl-7.15.5-9.el5.x86_64.rpm gnupg-1.4.5-14.x86_64.rpm rhnlib-2.5.22-3.el5.noarch.rpm rhpl-0.194.1-1.0.2.x86_64.rpm rpm-python-4.4.2.3-18.el5.x86_64.rpm up2date-5.10.1-41.8.el5.x86_64.rpm

After it:
[root@ed01db01 Server]# rpm -Uhv --nodeps curl-7.15.5-9.el5.x86_64.rpm gnupg-1.4.5-14.x86_64.rpm rhnlib-2.5.22-3.el5.noarch.rpm rhpl-0.194.1-1.0.2.x86_64.rpm rpm-python-4.4.2.3-18.el5.x86_64.rpm up2date-5.10.1-41.8.el5.x86_64.rpm
warning: curl-7.15.5-9.el5.x86_64.rpm: Header V3 DSA signature: NOKEY, key ID 1e5e0159
Preparing...                ########################################### [100%]
        package rpm-python-4.4.2.3-18.el5.x86_64 is already installed


But up2date have not been installed after this cmd!
One modification was required to finally set up up2date:

[root@ed01db01 Server]# rpm -Uhv --nodeps gnupg-1.4.5-14.x86_64.rpm rhnlib-2.5.22-3.el5.noarch.rpm rhpl-0.194.1-1.0.2.x86_64.rpm  up2date-5.10.1-41.8.el5.x86_64.rpm
warning: gnupg-1.4.5-14.x86_64.rpm: Header V3 DSA signature: NOKEY, key ID 1e5e0159
Preparing...                ########################################### [100%]
   1:rhpl                   ########################################### [ 25%]
   2:gnupg                  ########################################### [ 50%]
   3:rhnlib                 ########################################### [ 75%]
   4:up2date                ########################################### [100%]
[root@ed01db01 Server]#


Only ater such long way we have working up2date !

And the key file appeared:
[root@ed01db02 Server]# ls -l /usr/share/rhn/RPM-GPG-KEY
-rw-r--r-- 1 root root 1397 Nov 11  2007 /usr/share/rhn/RPM-GPG-KEY

And next step ewre executed well:
[root@ed01db01 Server]# rpm --import /usr/share/rhn/RPM-GPG-KEY
[root@ed01db01 Server]#

And the up2date --register works well.

Tuesday, May 15, 2012

Weak password: not enough different characters or classes.

Для нового класса студентов понадобилось установить какой-нибудь простенький пароль на cellmonitor. Однако возникла Problem:

[root@ed01cel02 cellos]# passwd cellmonitor
Changing password for user cellmonitor.

You can now choose the new password or passphrase.

A good password should be a mix of upper and lower case letters,
digits, and other characters.  You can use a 5 character long
password.

A passphrase should be of at least 3 words, 5 to 40 characters
long and contain enough different characters.

Alternatively, if noone else can see your terminal now, you can
pick this as your password: "please_goal!Burma".

Enter new password:
Weak password: not enough different characters or classes.
Try again.

Solution: change enforce=everyone ->  enforce=none

[root@ed01cel02 cellos]# cd /etc/pam.d/

[root@ed01cel02 pam.d]# cat system-auth
#%PAM-1.0
# This file is auto-generated.
# User changes will be destroyed the next time authconfig is run.
auth        required      pam_env.so
auth        required    pam_unix.so try_first_pass nullok
#auth        required      pam_deny.so

account     required      pam_unix.so

password    requisite     pam_passwdqc.so min=5,5,5,5,5 similar=deny enforce=everyone max=40
password    sufficient    pam_unix.so try_first_pass use_authtok nullok md5 shadow remember=10
password    required      pam_deny.so

session     optional      pam_keyinit.so revoke
session     required      pam_limits.so
session     [success=1 default=ignore] pam_succeed_if.so service in crond quiet use_uid
session     required      pam_unix.so

End think about return enforce=everyone back after new password is set

Saturday, April 28, 2012

Oracle Unbreakable Enterprise Kernel Release 2

Попался мне интересный документ: Oracle Unbreakable Enterprise Kernel Release 2 Release Notes

Весьма рекомендую всем почитать.
Что привлекло мое внимание в новом ядре :
Transparent Huge Pages - dynamically allocating hugepages
Memory compaction - merge used pages into a new big block of contiguous pages
Transmit Packet Steering - spreading of outcoming network traffic across all CPUs in system


  

Wednesday, January 25, 2012

Linux IO scheduler

Linux IO scheduling controls the algorithm for processing disk IO requests in appropriate order. Choosing adequate algorithm may have big impact on the IO performance. IO scheduling is determined by the Linux kernel. From the 2.6 there are four IO schedulers:
  • Completely Fair Queuing (CFQ)
  • Deadline
  • NOOP
  • Anticipatory
The scheduler is determined by elevator option in the /boot/grub/grub.conf file:
elevator= cfq | deadline | noop | as

With 2.6 kernels, it is also possible to change the scheduler for particular devices
during runtime:
# echo deadline > /sys/block/sdb/queue/scheduler

 Steve Shaw and Martin Bach writes:
"CFQ is the default scheduler for Oracle Enterprise Linux. It balances IO requests across all available resources. The Deadline scheduler attempts to minimize the latency of IO requests with a round robin–based algorithm for real-time performance. The Deadline scheduler is often considered more applicable in a data warehouse environment, where the IO profile is biased toward sequential reads. The NOOP scheduler minimizes host CPU utilization by implementing a FIFO queue, and it can be used where IO performance is optimized at the block-device level. The Anticipatory scheduler is used for aggregating IO requests where the external storage is known to be slow, but at the cost of latency for individual IO requests. "

But in the 6.2 the default  scheduler in was changed:
"Default IO scheduler

    For the Unbreakable Enterprise Kernel, the default IO scheduler is the 'deadline' scheduler.
    For the Red Hat Compatible Kernel, the default IO scheduler is the 'cfq' scheduler. "
 http://oss.oracle.com/ol6/docs/RELEASE-NOTES-U2-en.html

Tuesday, October 25, 2011

Huge Pages & Exadata

I decied to use Huge Pages in our Exadata and just opened /etc/sysctl.conf with vi but …

... vm.nr_hugepages was commented :
# bug 8268393 remove vm.nr_hugepages = 2048

I Пришлось пойти на металинк. Нота относится к Экзадате на НР:
Вот о чем речь:
-----------------------------------------------------
Disable Hugepages on the Database Servers
The current database image allocates 4G of hugepages that in many cases do not get used for various reasons. Some common examples of why the hugepages aren't used are:
  • Allocation requested at SGA creation time is larger than the hugepage config
  • Automatic Memory Management is being used which cannot leverage hugepages
  • The oracle user cannot lock the requested amount of memory because the memlock limit is either not specified or under-configured in /etc/security/limits.conf
Considering these issues, and the lack of benefit for a Data Warehouse environment, it is recommended that hugepages be disabled on the database servers.
Please note that hugepages should   NEVER   be disabled on the Exadata cells.
Bug 8268393 has been filed to make this change permanent in the database side image, and includes steps to do this manually in the work-around section.
-----------------------------------------------------

В общем:
-  упреки от пользователей, нежелающих разобраться почему не стартует их БД +
- несовместимость  АММ с большими страницами
заставили Оракл отказаться от этого функционала и Оракл даже специально оформил якобы баг, чтобы запретить большие страницы на Экзадате.

Жалко, это очень полезный функционал. Особенно для больших систем.

В чем-то Оракл здесь неправ. Большие страницы приносят свою пользу в больших системах.
В общем, если включить HugePages, то ошибок не будет.

# ocrconfig -add +DATA PROT-30: The Oracle Cluster Registry location to be added is not usable. PROC-50: The Oracle Cluster Registry locatio...