Sunday, April 27, 2014

Exadata Documentation

Exadata Documentation can also be found on My Oracle Support.
It comes with Patch 10386736.




DOCUMENTATION FOR EXA 11.2 & 12.1
Last Updated 19-Apr-2014 03:00 (8 days ago)
Product Oracle Exadata Storage Server Software
Release Oracle SAGE 12.1.1.1.0
You will find following guides and information when you unzip the downloaded patch file.

Oracle Exadata Database Machine Documentation 

Exadata Database Machine Owner's Guide
Exadata Storage Server Software Release Notes
Exadata Storage Server Software User's Guide
Auto Service Request Quick Installation Guide for Oracle Exadata Database Machine
Enterprise Manager Exadata Management Getting Started Guide
Exadata Database Machine Security Guide
Database Licensing Information
Exadata Database Machine Licensing Information
Exadata Database Machine Safety and Compliance Guide

Saturday, April 26, 2014

Exadata -- For Exadata Database Machine Admins -- Filtered Information


This post will be based on the information concerning Exadata Administrators..
You will find information gathered on the way for being certified on Exadata.. 
The information provided will be item by item, and I think this post will be useful for Exadata Admins even though they already have field experience.. 

This post will be unstructured,.I mean the information provided in this document is not grouped into subtitles, but it will present the whole picture about Exadata Administration, when you will read the all of it.

The information provided in this document will be like snips from the Exadata knowledge..
In order to understand this documents, you should already know the basic concepts of Exadata.( like Storage indexes, smart scan and etc.. )

Let start;

You can make indexes invisible to increase the chance for using storage indexes.
You can load your data stored on the filtered columnd to maximize the benefit of storage indexes.
Bind variables can be used with storage indexes.
Oracle Database QOS management can offer recommendations for the CPU bottlenecks. QOS management can not provide recommendations for the Global Cache resource bottlenecks. OQS management cannot resolve IO resource bottlenecks, too.
For media based backups, it s recommended to allocate equivalent number of channels and instances per tape drive.. For this type of backups, the network cables between Exadata and media servers should be connected through Exadata Database nodes. 
Bonding on Exadata can be used both for reliability and load balancing, but when I look to a Production Exadata Machine , which is X3, I see that the bonding is configured for reliability.

cat /proc/net/bonding/bondeth0
Ethernet Channel Bonding Driver: v3.2.3 (December 6, 2007)
Bonding Mode: fault-tolerance (active-backup)
....
BONDING_OPTS="mode=active-backup miimon=100 downdelay=5000 updelay=5000 num_grat_arp=100"


active-backup or 1 -> Sets an active-backup policy for fault tolerance.
Transmissions are received and sent out via the first available bonded slave interface.
Another bonded slave interface is only used if the active bonded slave interface fails.


To guarentee proper cooling for an Exadata Machine, perforated floor tiles should be placed at the front, because the air flow is from front to back.

Creating multiple grid disks on a single disk in Exadata provides us multiple storage pools with different performance characteristics and multiple pools that can be assigned to different databases.

Here is a general information about Disk layout in Exadata
Physical disk -> Lun -> Cell Disk( its like a filsystem on Lun..) -> GridDisk ( it s like a partition)
Lastly Grid disks are served to the ASM diskgroups..
Note that , we can create flash based ASM diskgroups,too. They occupy may share space with the Flash Cache on the flash disks.

Things to consider for Exadata migrations,
For Transportable Database method-> Source database should be 10.2.0.5 and little endian
For Data Pump method ->, 10.2.0.5 is good, but it is time consuming..
For Data guard physical method ->, Source platform must be Linux, Solaris x86 or Windows and source must be on 11.2.
For ASM rebalance method -> 11.2 Linux x86 64 database that uses  ASM w/ 4MB AU.
Transportable tablespaces method -> Big endian source >= 10.1 or Little endian source>=10.1, <11.2  is needed
Logical standby method -> Source does not need to be on 11g. Logical Standby is not supported from HPUX to Linux migrations.

Also here is an useful information about the methods:

Standby Physical Standby:
-------------------------------------------------
Source platform must be Linux, Solaris x86 or Windows (see Document 413484.1) and source on 11.2
Source database in ARCHIVELOG mode
NOLOGGING operations to the database are not permitted, i.e. FORCE LOGGING is turned on at the database level

Transportable Database
-------------------------------------------------

Source system must be 10gR2 (10.2.0.5) and little endian
Stricter service level requirements that obviate the required downtime with Oracle Data Pump

Transportable Tablespaces
-------------------------------------------------
Any source platform
EBS Release 12 with source database RDBMS release 10gR2 (10.2.0.5) or higher
EBS 11i with source database RDBMS release 11.2
Service levels cannot tolerate a significant outage.  Using transportable tablespaces instead of Oracle Data Pump export/import can significantly reduce the outage time for large (> 300 GB) EBS databases.  Thorough testing will determine the precise outage time.
For a point of reference, tests on an Oracle Exadata Database Machine quarter rack with the Rapid Install Vision database (about 300 GB) took about 12 hours.  This time should remain about the same regardless of the amount of data in the database.  This is because the metadata creation takes the longest time in the migration process and accounts for the bulk of time.

Oracle Data Pump
-------------------------------------------------
Any source platform
Source database is RDBMS release 10.2 or higher
To implement Oracle Exadata Database Machine best practices on the target
Service levels can tolerate a significant outage.
For a point of reference, tests on an Oracle Exadata Database Machine quarter rack with the Rapid Install Vision database (about 300 GB) took about 24 hours (export - 7:42; import - 16:42) using Network Storage and no dump file copy (i.e. the export dump storage was mounted on the source and the target).
Timings will vary depending on your system configuration and increase as the amount of data increases.

In Exadata, Oracle Enterprise Agents must be deployed to the compute nodes. .Oracle Exadata Plug-in deployed with the Agent. Plugins allow you to monitor the following key components of Exadata machine. There are several plugins for Grid Control and Cloud Control (such as Avocent MergePoint Unity Switch, Cisco switch, Oracle ILOM, Infiniband switch, PDU) .. Note that : A trap forwarder is required to catch cisco switch and kvm traps due to a port mismatch..
Agent communicates with Storage Server and Infiniband Switch targets directly. Oracle Exadata Plug-in also monitors the other DBM components. Oracle Enterprise Manager 12c agent collects data and communicates with the remote Enterprise Manager Repository.

In Exadata X3, we have 512 MB flashlogs on the storage servers.
In compute nodes, we have raid 5 arrays;

Virtual Drives
Virtual drive : Target Id 0 ,VD name DBSYS
Size : 556.929 GB 
State : Optimal
RAID Level : 5

Exadata -> Action plan to replace a Flash disk when Griddisks are created them,
Ref: Replacing FlashCards or FDOM's when Griddisks are created on FlashDisk's (Doc ID 1545103.1)

If the flash card needs to be replaced, drop the disks used in +ASM for the FlashDisk
Drop FlashCache / Flashlog and delete celldisks of type Flashdisk
Shutdown
Replace
create flashdisks
Now create your griddisk's back on DiskType: FlashDisk
add back the disks to the diskgroup. (it is optional ->If you used 'force' when dropping the disks from ASM then Exadata auto management should automatically add these disks back into ASM.)

Note that, we have different failure types in Exadata. For example a disk which is in predictive failure state must be replaced immediately. ASM will drop these kind of disks automatically from the associated disk group , and start a rebalance operation.
For flash a flash disk also; if the disk is in predictive failure then ->
If the flash disk is used for flash cache, then flash cache will be disabled on this disk thus reducing the effective flash cache size. If the flash disk is used for flash log, then flash log will be disabled on this disk thus reducing the effective flash log size. If the flash disk is used for grid disks, then Oracle ASM rebalance will automatically restore the data redundancy..

If you want to change a memory dimm or want to make a similar activity, you just need to shutdown the affected cell.. The database, and the other cell servers will not be affected from this operation.
What you have to do is;
Ensure asmdeactivityoutcome of the Griddisks first..
Then inactive all the griddisks on that cell and ,shutdown the cell using shutdown -h now ..

quartery battery learn cycle is a normal maintanence activity.. It is used for charging the batteries. When this kind of activity is happening, you can end up with a performance degregation in terms of I/O, as this event may require the related device to be put in the write through mode.. Keep that in mind and, If you can or need, set the device back to write back mode...

Configure and use smart flash logs if you need a better LGWR performance, as LGWR writes redo data both flash and disk in parallel.. It considers whichever of these writes completes first, as done..
Smart Flash Logs are a new feature comes with 11.2.2.2.4 cell software. They are not for reading, they are used like a circular buffer for Redo writes.. Smart Flash logs can enhance the performance of an OLTP database. By default they occupy 512 MB per cell. (32 MB size per flash disk *16 flash disk) space in the Flash Cards.. So Smart Flash logs reduce the size of Flash Cache in a manner.
Note that , Flash Smart logs can be enabled or disabled using IORM according to the databases if needed..
Also LGWR writes and Controlfile IO will be automatically on high priority in IORM..  On the other hand, DBWR IO is managed automatically at normal priority.
Flash logs if needed somehow, can be dropped using drop flashlog command through the cellcli utility residing on the storage servers..

Consider using CTAS for bulk data loading from external tables to Exadata.. CTAS automatically uses direct path load.. Insert /*+append*/ can also be used for this kind of data loading, as by using the append hint, oracle will use direct path loading in insert operations, too..

cellip.ora is the configuration file , if you want to separete cells which are connected by the asm instances. This file is located on where ASM resides(compute node), and it basically tells ASM which cells are available. cellip.ora is located in every compute node, and its contents are like following;

cat /etc/oracle/cell/network-config/cellip.ora
cell="192.168.10.10"
cell="192.168.10.11"
cell="192.168.10.12" 

This file should also be used if you want to add an expansion rack to the storage grid.

In default configuration , DATA and RECO diskgroups are build on top of non-interleaving disks. For detailed information about Interleaving , please see my following post http://ermanarslan.blogspot.com.tr/2013/12/exadata-zbr-and-interleaving.html
Also, in default configuration, we dont have any free space in Flash Disks;

CellCLI> list celldisk where name='FD_15_cel01' detail
name: FD_15_cel01
comment:
creationTime: 2012-01-10T10:13:06+00:00
deviceName: /dev/sdy
devicePartition: /dev/sdy
diskType: FlashDisk
errorCount: 0
freeSpace: 0
id: 8ddbd2c8-8446-4735-8948-d8aea5744b35
interleaving: none
lun: 5_3
size: 22.875G
status: normal

Note that , you can use infiniband to connect an Exadata to a Exalogic. You can also use infiniband to connect and Exadata to a Oracle ZFS Storage ZS3, as ZS3 has a native infiniband connectivity.
Alternatively, Sun Zfs Storage 7420 Appliance can be connected to Exadata directly from infiniband.. 
In addition, you can connect any media servers which have infiniband cards, to Exadata via infiniband. Then you can connect those media servers to tape libraries to maximize the tape backup thorughput..
In terms of tape backups, Oracle docs suggest to have Disk-to-disk-to-tape, in other words D2D2T strategy, which allows keeping old backups on tape while retaining new/fresh backups on disks for achieving fast recovery times.  You can consider the following Oracle Slide as a good Exadata-tape backup scenario;


Uncommited transactions and migrated rows can cause cell single block physical reads even if you are doing a Full table scan. Note that: Single block reads and Multi block reads may benefit from the Flash Cache,
Also note that, smart scan can not be done against Index Organized Tables and clustered tables.

When you need to apply a bundle patch  to Oracle Homes in Exadata, you will need to use oplan utility.
oplan utility generates instructions for applying patches, as well as instructions for rollback. It generates instructions for all the nodes in the cluster. Note that, oplan does not support DataGuard .. Oplan is supported since release 11.2.0.2 . It basically eases the patching process, because without it you need to read Readme files and extract your instructions yourself..
It is used as follows;
as Oracle software owner,(Grid or RDBMS) execute oplan;
$ORACLE_HOME/oplan/oplan generateApplySteps <bundle patch location>
it will create you patch instructions in html and txt formats;
$ORACLE_HOME/cfgtoollogs/oplan/<TimeStamp>/InstallInstructions.html
$ORACLE_HOME/cfgtoollogs/oplan/<TimeStamp>/InstallInstructions.txt
Then, choose the apply strategy according to your needs and follow the patching instructions to apply the patch to the target.
That 's it.. 
If you want to rollback the patch;
execute the following;(replacing bundle patch location)
$ORACLE_HOME/oplan/oplan generateRollbackSteps <bundle patch location>
Again,   choose the rollback strategy according to your needs and follow the patching instructions to rollback the patch from target.

It is mandatory to know which components/tools are running on which servers on Exadata;
Here is the list;

DCLI -> storage cell and compute nodes, execute cellcli commands on multiple storageservers
ASM -> compute nodes -- it is ASM instance basically
RDBMS -> compute nodes -- it is Databas software
MS-> storage cells, provides a Java interface to the CellCLI command line interface, as well as providing an interface for Enterprise Manager plugins.
RS -> storage cells,RS, is a set of processes responsible for managing and restarting other processes.
Cellcli -> storage cell, to run storage commands
Cellsrv ->storage cell. It receives and unpacks iDB messages transmitted over the InfiniBand interconnect and examines metadata contained in the messages
Diskmon -> compute node, In Exadata, the diskmon is responsible for:,Handling of storage cell failures and I/O fencing
,Monitoring of Exadata Server state on all storage cells in the cluster (heartbeat),Broadcasting intra database IORM (I/O Resource Manager) plans from databases to storage cells,Monitoring or the control messages from database and ASM instances to storage cells ,Communicating with other diskmons in the cluster.


The following strategy should be used for applying patches in Exadata:

Review the patch README file (know what you are doing)
Run Exachk utility before the patch application (check the system , know current the situation of the system)
Automate the patch application process (automate it for being fast and to minimize problems)
Apply the patch
Run Exachk utility again -- after the patch application
Verify the patch ( Does it fix the problem or does it supply the related stuff)
Check the performance of the system(are there any abnormal performance decrease)
Test the failback procedure (as you may need to fail back in Production, who knows)


Multiple Grid disks can be created on a single Cell Disks, as you know. While creating multiple Grid disks on a Single Disk, you can end up having multiple disks which have different performance characteristics . In order to have more balanced disk layout, you can use interleaving options or Intelligent Data Placement technology which is based on ASM..

The internal Infiband network used in Exadata transmits IDB messages between compute nodes and storage servers as well as Rac interconnect traffic packets between the compute nodes in the cluster.

Dbfs or a Nfs filesystem can be used as a stage for loading data from the external tables. If choosing Nfs to load data, the Nfs share should be mounted to the preferred compute node..
Dbfs can be created on DBFS_DG , as well as on a standart filesystem..
It can enhance performance if you need to bulk load data in to your Production System , which resides on Exadata.. By using DBFS created on ASM, you will have a parallelization in storage, which will enhance perfomance of IO.

Diagget.sh script can be used to gather the diagnostic information with software logs, trace files and OS information..

Exadata storage server has alerts defined on it by default.. This alerts are based on some defined metrics.. We can define new metrics also, this metrics will persist accross cell node reboots.

If you have a 11.1.0.2 database(little endian) and want to migrate it to Exadata with minimum downtime, you can upgrade it to 11.2.0 and use Data pump physical to minimize downtime or alternatively, you can use Golden Gate for this. Ofcourse you can use datapump, as well, but your with datapump the downtime will be significant..

According to the Non-interleaving disk configuration in Exadata(which is default), we can say that the first Grid diskdisks created using Create Griddisk... command will have the best performance.  So the Diskgroup created based on the first created griddisk will have better performance than the other Diskgroup on the Same Exadata Machine. ...

If you want to do  some administrating stuff on Exadata Storage servers, like dropping the cell disk and etc;
you can use celladmin OS account in storage servers to do this kind of an operation. You can use cellcli(on every cell) or dcli(from one cell to administer all of the cells) utility to execute your commands..
DCLI is a pyhton script. It is used to execute command on the cells remotely without login in to them.. (In first execution it create ssh keys with the following -> dcli -k -g mycells.)

Note that, we have firewall (iptables) configured with Oracle Supplied rules on cell servers.

In OLTP systems,
Exadata write-back flash cache can enhance performance , as it also caches database block writes.
Also Flash log can be useful for enhanching perfomance of OLTP systems, as fast response time for Log Writer is crucial in OLTP..
For Big OLTP systems, Flash cache and Flash log can enhance performance and High Capacity Disks in Exadata can meet the storage size needs..

Note that , IDP, IDB and IORM manager can only seen on an Exadata Environment.

IORM is used for managing the Storage IO resources.

Here is a diagram for the description of the architecture of IORM. (reference: Centroid)


So we can manage our IO resources based on the Categories, Databases and consumer groups. There is a hierarchy as you see in the picture.. The hierarchy used to distribute I/O.
IORM must be used on Exadata if you have a lot of databases running on Exadata Machine.. IORM is a friend of consolidation projects, in my opinion..

If you configured ASR manager in your environment, note that : database uses SNMP to transfer the notifications from database to ASR Manager and these notifications are forwarded using SNMP from ASR to Enterprise Managern In addition Faults are tranferred to the Oracle securely using https.

The fault telemetry sent from ASR manager is in the below format;

Telemetry Data:

System-id: System serial number.
¦ Host-id: hostname of the system sending the message.
¦ Message-time: time the message was generated.
¦ Product-name: Model name of the server


As known, Exachk is the utility to validate an Exadata Environments. We check the system using this utility time to time(before patches, after patches and etc..) Also, it can be scheduled to run regularly.. To schedule exachk, we can create a cron job or we can create a job in Enterprise Manager..

Compression on Exadata can only be done on the Compute nodes.. Decompression, on the other hand, can be done on Compute Nodes or Storage Cells.. Decompression can be done on Storage Cells if the associated operation is based on the Smart Scan.. Same rule applies for Encryption and Decryption, too..

Here is a general information about compression types:
BASIC compression, introduced in Oracle 8 already and only recommended for Data Warehouse
OLTP compression, introduced in Oracle 11 and recommended for OLTP Databases as well
QUERY LOW compression (Exadata only), recommended for Data Warehouse with Load Time as a critical factor  --> HCC
QUERY HIGH compression (Exadata only), recommended for Data Warehouse with focus on Space Saving --> HCC
ARCHIVE LOW compression (Exadata only), recommended for Archival Data with Load Time as a critical factor --> HCC
ARCHIVE HIGH compression (Exadata only), recommended for Archival Data with maximum Space Saving --> HCC

If your index creation takes long time in Exadata, consider the following information:
Cell single block physical read can impact the peformance of an index creation activity. Cell single block physical reads are like db file sequential reads in a tradional system.
Migrated and chained rows can cause cell single block read events on an Exadata Machine.  Also uncommited rows during a query is handled based on the consistency.. For supplying the consistency, Database nodes may require additional blocks. These blocks are sent by the Cell servers to the database nodes.. This activity can cause cell single block physical reads , as well. If you have a lot of blocks in one of these conditions, then your index creation time can take long time..

When migrating to Exadata, you should consider Database Type(oltp or warehouse), size of the source database, the version of the source database and the Endian format of the Source Operating Systems.. By analyzing these inputs you can choose an optimal migration method and strategy.

In Exadata Enterprise Manager monitoring, the communication flow from the ILOM of the Storage Servers to Enterprise Manager is through the Storage Server's MSprocesses. ILOM sends data using snmp to MSprocess, and MSprocess send the data to the Enterprise Manager using snmp. Data is triggered and transfered through the snmp traps.
Based on the preset thresholds defined in ILOM, we can monitor motherboard,memory,power, and network cards of Database Nodes using Enterprise Manager. We can see the faults and alerts  produced for these hardware components from Enterprise Manager.

Oracle Auto Service Request (ASR) is a secure, scalable, customer-installable software feature of warranty and Oracle Support Services that provides auto-case generation when common hardware component faults occur. ASR manager can be installed on an external Oracle Linux or Oracle Solaris server. Also you can use one of Exadata db nodes for installing ASR manager (not preferred).
ASR manager communicates with Oracle using https..

In Exadata, some database work can be offloaded to Storage Servers as you know.. Besides operations like full table scan, single row functions and simple comparison operators, some joins can be offloaded to storage. Column filtering, Predicate filtering  and Virtual Column filtering are the things that can be offloaded to the Exadata Storage Servers, as well.
If you want to see all the functions that can benefit from smart scan; you can use the following sql:
select name from v$sqlfn_metadata
The output will be like;

>
<
>=
<=
=
!=
OPTTIS
OPTTUN
OPTTMI
OPTTAD
OPTTSU
OPTTMU
OPTTDI
OPTTNG
AVG
SUM
COUNT
MIN
MAX
OPTDESC
TO_NUMBER
TO_CHAR
NVL
CHARTOROWID
ROWIDTOCHAR
OPTTLK
OPTTNK
CONCAT
SUBSTR
LENGTH
INSTR
LOWER
UPPER
ASCII
CHR
SOUNDEX
ROUND
TRUNC
..
.....
.....

Operations like full table scan and fast full index scan executed in parallel always generate Smart scans.. For full table scan, you can see cell smart table scan, and for fast full index scan, you can see cell smart index scan events.. Fast full index scan operations can be executed through the smart scan because in fast full index scan, Oracle just reads the index block as they exist on the storage.. I mean a bulk read will be performed, and that makes Oracle to use smart scan even if it is an index operation.  Smart scan is performed during Direct path read operations.. Parallel queries use direct path read, that 's why they make use of smart scans.. 
So if we need to gather all together; in order to have smart scans, we need to execute queries in parallel, we need to use direct path reads towards the process memory and we need to have cell.smart_scan_capable parameter set to TRUE for our ASM diskgroups.

We have Sun servers in Exadata.. Compute nodes and storage servers are acutally Sun Servers. They have ILOM cards on them. These ILOM cardscan be used to administer these servers remotely.. For Example ILOM can be used to power-on database servers or open a remote console for a Storage Server.

Note that, if you want to check the status of all ports located in the Infiniband Switch, you can use Enterprise Manager or ibqueryerros.pl script located in the Infiniband switch.

To properly shutdown the Exadata, we need to first stop the database and grid services. We may use crsctl stop cluster -all command for this. Then we need to shutdown Exadata storage servers first. Then we need to shutdown Database servers. Lastly we need to power off the network switches and cut the power using power switches on PDUs. Why firstly shutdown storage servers? ( this is still a mystery for me)

After Exadata migrations, analysis are made to drop the indexes . This activity is done to increase the chance of using smart scans on our query, but care must be taken while dropping those indexes. For example, in an OLTP system we need fast response times for single block reads, so dropping an index may result a negative performance impact for some of our OLTP queries.. So it s better to drop an index after analyzing the queries of the corresponding application. You can use invisible indexes to see the difference .. You can check Execution plans to see if Oracle wants to use smart scan instead of an index access for a query.. According to your analysis, make the decision to drop the unnecessary indexes..

As you know Exadata contains Compute Nodes and Storage Nodes.. For Storage Nodes, you dont have the choice to use a different OS than Linux. Exadata Storage Servers are Linux, and will continue to be Linux. Oracle Linux servers are coming with Uek kernels.
 On the other hand, you can choose to have Solaris 11 rather than Linux for compute nodes. Operating systems are selectable at install time.

Note that, In Exadata we have different networks, like Management network and public network, infiniband network.

Following is a picture representing those networks; (it is for X4 actually, but it is useful)


Ssh, for example , works from the Management network, as it is an utility to manage the corresponding servers.  If you want to change the network that ssh listens from, you can use ssh config file to do that.

Following inputs can be used for configuring the Exadata Machine at install time..

Customer name
applcation
region
Timezone
Compute OS
Database Machine_Prefix
Admin Network -> Start Ip address for pool,pool size,Ending Ip address for pool,Subnet Mask,Gateway
Admin Network -> Database Server admin name, Storage Server Admin Name, ILON Name,
Client Network ->Start Ip address for pool,pool size,Ending Ip address for pool,Subnet Mask,Gateway, Adapter speed (1gbe/10gbe Base-t or 10gbE sFP+optical)
Infiniband NEtwork -> Start Ip address for pool, pool size, ending ip address for pool, subnet mask, compute node private name
Backup / Data Guard Ethernet NEtwork
Os configuration -> Domain , DNS, NTP , Grid ASM home os user ,asm dba group, asm home oper group, asm home admin group, rdbms home os user, rdbms dba group, oinstall group, rdbms home oper group, base loc for grid and rdbms --> you can set userid,groupid of all users and groups
Home and Database -> ınv location, grid home, db home loc, software install, db name, data reco group names and redundancy (cant change their sizes), block size , type DW or OLTP
Cell Alerting -> Enable Email Address
ASR conf
OCM conf
Grid Control agent

In order to actually have the Exadata Storage Servers send notifications via email (or alternately SNMP) each of the servers has to be configured with the appropriate settings. This is done using the ALTER CELL command in CELLCLI.

ALTER CELL smtpServer='mailserver', -
smtpFromAddr='exacel@blabla.com, -
smtpPwd='email_password', -
smtpToAddr='erm@blabla.com', -
notificationPolicy='critical,warning,clear', -
notificationMethod='mail'

The alerts may be stateful and stateless, as well..
If the alert is based on a threshold, it gets cleared automatically when it no longer violates the threshold;
For example; a filesystem alerts , as follows;

CellCLI> list alerthistory detail
name: 4_1
alertMessage: "File system "/" is 82% full, which is above the 80% threshold.
Accelerated space reclamation has started.
This alert will be cleared when file system "/" becomes less than 75% full."

Also alerts can be fired for Critical, Warning and Information purposes..

We have Storage indexes located in the physical memory of  the  Storage Servers.. Storage indexes are maintained automatically  the cellsrv and they are not persistent across reboots..(as they are in memory).. Cellsrv build these indexes based on the filter columns of the offloaded queries.
In Storage indexes, Oracle builds range regions based on the column values. By the use of storage indexes Oracle can easily find where to look in the storage for a given column value..
Maximum 8 columns for a table are indexed per storage region, but different storage regions can have different columns indexed for the same table.

Storage servers are very sensitive environments in Exadata, so Oracle doesnt support a lot of activities on them.. For example, we can change root password of these servers or we can set up a ssh equivalence for cellmonitor users(http://docs.oracle.com/cd/E11857_01/install.111/e12651/pisag.htm#CIHJGEHI)
 But for example, if we want to upgrade the storage server software, we need to use patchmgr utility..
the patchmgr utility is a tool Exadata Database Administrators use to apply (or rollback) an update to the Oracle Exadata Storage Cell.
Here is another restrictions is ;
Oracle Exadata Storage Server Software and the operating systems cannot be modified, and customers cannot install any additional software or agents on the Exadata Storage Servers.

Lastly, I will mention about placing multiple Exadata Machines in a System Room/Data Center.
If you have multiple Exadata Database machines, you need to place them side by side while ensuring the exhaust air of one rack does not enter the airinlet of another. If you have multiple clusters running on several Exadata Machines, you can place the racks that are part of a common cluster together/side by side..


That' all for now. I hope you will find it useful.

Thursday, April 24, 2014

RDBMS-- Oracle error -449: ORA-00449: background process 'MMON' unexpectedly terminated with error 448

You can encounter this error on EBS 12.1 and above , especially in the plsql 's that are used for authentication purposes.
For example changing an application users' password or in the login phase of Oracle Applications..
Produced error will be something like below;

Oracle error -449: ORA-00449: background process 'MMON' unexpectedly terminated with error 448 has been detected in FND_WEB_SEC.VALIDATE_PASSWORD

The definitions of the errors are as follows;

oerr ora 448
00448, 00000, "normal completion of background process"
// *Cause:  One of the background processes completed normally (i.e. exited).
//         The background process thinks that somebody asked it to exit.
// *Action: Warm start the system.
 oerr ora 449
00449, 00000, "background process '%s' unexpectedly terminated with error %s"
// *Cause:  A foreground process needing service from a background
//          process has discovered the process died.
// *Action: Consult the error code, and the trace file for the process.

So, an error 449 is produced because of the error 448, which means our foreground process(LOCAL=NO) needs MMON but MMON background process acutally had been exited already.

Also when this error is produced,  if you check the existence of mmon process , you will see that it s not there.. (ps -ef |grep mmon)

Best way to fix this error is to restart application and database services, as suggested by Oracle Support.
Also MMON trace can be gathered to check out the root cause --if necessary, but Oracle Support says that this problem can be faced in EBS 12.1 e nvironments, and it is because of a process lock , so restarting app+db services should be enough to fix it..


Tuesday, April 22, 2014

Linux -- understanding the "free" command output, reclaimable vs free

Nowadays, we have processes running on Linux .. They allocate high amount of memory regions and sometimes we suffer from this high memory consumptions.
On the other hand, by just looking to output that are generated by tools like Top , is not enough to make a decision about the memory consumptions..
Nowadays, I have seen that even the developers use free command to analyze the memory consumption of the OS that their applications are running on.. I also have seen that sometimes the output of free command is interpreted wrongly, and most of the time it has been thought that the system is suffering from the memory.
That's why I m writing this post , and sharing output of free -m command in an understandable format that summarizes the information that can be used to analyze the free command output.
Dont forget that the memory is a resource that is meant to be used %100 and Linux memory concept is designed for efficiency.

 Example of free -m command output , in Linux;

[root@nagios ~]# free -m

                      Total                       Used                  Free             Shared               Buffers          Cached

Mem:              512                         392                    119                 0                           28             249
                       Total Phys             Used Phys             Free Phys      Shared mem                 Buffers    Fast access                                   Mem                 Mem               Mem         across procs            

-/+ buffers/cache:                             113                    398
                                                Actual Used           Actual Free 
                                                  Phys Mem              Phys Mem

Swap:               2975                         0                           2975
                  Total Swap               Used Swap              Free Swap

So, the memory for buffers and cached memory can reclaimed if needed. So they can be thought that they are free..

If we need to calculate the used memory and free memory we should use the following formula;

Used Memory = Used Phys - Buffers - Cached -> in this example  392 - 28 - 249 = 115
Free Memory  = Free Phys + Buffers + Cached -> in this example 119 + 28 + 249 = 396 

Look for the -/+ buffer cache lines now , seems familiar :) Used 113 (almost 115) , Free 398 (almost 396)

So if you use free -o command you wont have a line starting with -/+ buffers/cache 
free -o -m

                    total       used       free     shared    buffers     cached
Mem:           512        401        110          0         34        252
Swap:         2975          0       2975

So use "-o" option  if you want to calculate the actual used and free memory yourself :)
The -o switch disables the display of a "buffer adjusted" line.  If the -o option is not specified, free subtracts buffer memory from the used memory and adds it to
       the free memory reported.

Note that, Linux is designed to use swap, so if you some used Swap in free command output, dont panic. Analyze swap in/out operations using "sar" or a similar tool..

Also note that: If you want to drop those caches , you can use /proc/sys/vm/drop_cache.
First issue sync command (for dirty things) , execute echo 3 > /proc/sys/vm/drop_caches

Sometimes we use 3 sync commands repeatedly.. Do we need 3 syncs? That is another story :)


Another place to look for memory consumption is  /proc/meminfo

Here is an example meminfo:

MemTotal: 524288 kB
MemFree: 109992 kB
Buffers: 37736 kB
Cached: 258604 kB
SwapCached: 0 kB
Active: 141500 kB
Inactive: 203412 kB
HighTotal: 0 kB
HighFree: 0 kB
LowTotal: 524288 kB
LowFree: 109992 kB
SwapTotal: 3047416 kB
SwapFree: 3047416 kB
Dirty: 340 kB
Writeback: 0 kB
AnonPages: 48564 kB
Mapped: 13780 kB
Slab: 20736 kB
PageTables: 9344 kB
NFS_Unstable: 0 kB
Bounce: 0 kB
CommitLimit: 3309560 kB
Committed_AS: 192008 kB
VmallocTotal: 34359738367 kB
VmallocUsed: 3600 kB
VmallocChunk: 34359734759 kB

As you see , it has similar values in terms of memtotal, memfree , buffers and cache
Okay we said that , the actual free memory is free + cache+ buffers, but we also know that cache and buffers are not free actually, on the other hand; they are reclaimable.

So is the free memory and reclaimable memory the same thing ?
In my opinion , they are not the same thing in theory, actually.
We can put anything to our free pages, but our file data and metadata are put on the reclaimable memory.
In other words, in a healthy system, reclaimable memory should be non-zero .. There should be room to grow in memory without effecting those caches.
That is , maybe we can meet the user process requirements by reclaiming the pages from the caches, but this may also effect negatively on performance in terms of IO , because we are freeing the cache ..(It is been said on some articles, that in the last 1MB of the freeing caches, there will be a noticable performance degregation in a crowded system.) 

Thursday, April 17, 2014

OpenSSL- Security Bug from Oracle 's Perspective

In April 2014, a security bug came to light in Open SSL cryptographic software library.
This bug is named as Hearthbleed, and it makes it possible to steal sensitive information when using the secure connection to the applications through ssl.
Information about this bug can be gathered by the following web site: http://heartbleed.com

Anyways, as our focus is on Oracle , it is good to know the affected and unaffected product of Oracle in the first place.. So I m sharing the affected and unaffected Poduct lists supplied by Oracle..
Note that, I m happy to say that EBS 11i or R12, which are my main focuses, seem not affected from this security bug.. On the other hand Oracle Linux 6 seems affected, and that's why has to be patched .. Goldengate also is still in investigation process..

Reference: Oracle

Not affected Products:

  • Advanced Lights Out Manager (ALOM) [Product ID 9843/ALOM/ALOM]
  • ALOM-CMT [Product ID 9846/SYSFW-ALL/ALOM-CMT]
  • Audit Vault [Product ID 1977,9749]
  • Brocade(McData) Fiber Channel Switches and Management Software [Product ID 9864]
  • Cisco MDS Fiber Channel Switches and Management Software [Product ID 9865]
  • E-Business Suite 11i
  • Enterprise Manager Cloud Control
  • Enterprise Manager Grid Control [Product ID 1370]
  • Enterprise Manager Grid Control Plug-ins and Connectors
  • Enterprise Manager Ops Center
  • Exadata [Product ID 2546]
  • Exalogic
  • FLEXCUBE Lending and Leasing 12.0, 12.1, 12.5 [Product ID 10484]
  • Hyperion BI
  • Hyperion Essbase [Product ID 4379]
  • ILOM [Product ID 9849]
  • JDE EnterpriseOne Tools [Product ID 4781]
  • MySQL Enterprise Backup [Product ID 4629]
  • NM2 IB switches [Product ID 10140]
  • NM2-36P InfiniBand switches [Product ID 10140]
  • Oracle Access Manager 10g and 11g Webgates [Product ID 5565]
  • Oracle Access Manager 10g Server [Product ID 5565]
  • Oracle Agile Engineering Data Management [Product ID 4436]
  • Oracle API Gateway 11.1.1 and 11.1.2 [Product ID 9195]
  • Oracle Business Intelligence Enterprise Edition [Product ID 2025]
  • Oracle Commerce Guided Search / Oracle Commerce Experience Manager [Product ID 9633/MDEX]
  • Oracle Communications ASAP [Product ID 2260]
  • Oracle Communications Billing and Revenue Management [Product ID 2136]
  • Oracle Communications Border Gateway
  • Oracle Communications Core Session Manager
  • Oracle Communications EAGLE Application Processor Query Server 15.0, 15.0.2 [Product ID 11117]
  • Oracle Communications EAGLE LNP Application Processor 10.0 [Product ID 11118]
  • Oracle Communications Enterprise Communications Broker
  • Oracle Communications Enterprise Session Border Controller
  • Oracle Communications IP Service Activator [Product ID 2261]
  • Oracle Communications Order and Service Management [Product ID 2270]
  • Oracle Communications Performance Intelligence Center 9.0.x [Product ID 11044]
  • Oracle Communications Policy Management 9.3, 9.4, 9.7, 9.8, 10.4, 11.0 [Product ID None yet]
  • Oracle Communications Security Gateway
  • Oracle Communications Service Broker Engineered System Edition [Product ID 9056]
  • Oracle Communications Session Router
  • Oracle Communications Session Tunneled Session Controller
  • Oracle Communications Session Tunneled Session Controller SDK
  • Oracle Communications Subscriber Data Management 9.1, 9.2 , 9.3 [Product ID None yet]
  • Oracle Communications Subscriber Profile Repository 9.0
  • Oracle Communications Subscriber-Aware Load Balancer
  • Oracle Communications Unified Session Manager
  • Oracle Database Appliance Software [Product ID 9435]
  • Oracle Database Firewall [Product ID 8958,9749]
  • Oracle DayBreak [Product ID 9696]
  • Oracle Eagle LNP Provision System 10.0.0
  • Oracle Endeca Server [Product ID 10217]
  • Oracle GlassFish Server 3.x.x [Product ID 8493]
  • Oracle Key Manager [Product ID 10052]
  • Oracle Linux 5 [Product ID 1309]
  • Oracle Real-time Scheduler [Product ID 2238]
  • Oracle Secure Backup 10.3, 10.4 [Product ID 1522]
  • Oracle Secure Global Desktop 4.x, 5.x [Product ID 8539]
  • Oracle Switch ES1-24 [Product ID 9889/OPUS24]
  • Oracle System Assistant [Product ID 10015]
  • Oracle Transportation Management 6.0, 6.1, 6.2 [Product ID 1991]
  • Oracle Tuxedo [Product ID 5433]
  • Oracle Virtual Desktop Infrastructure 3.3 to 3.5 [Product ID 8540]
  • Oracle VM [Product ID 4455]
  • Oracle VM VirtualBox 4.2, 4.3 [Product ID 8370]
  • Oracle WebLogic Web Server Plug-In 1.0 [Product ID 5242/PLUGIN]
  • Oracle ZFS Storage Software [Product ID 10026]
  • PeopleSoft Application Products including Campus Solutions [Product ID 5085]
  • Qlogic Fiber Channel Switches and Managment Software [Product ID 9866]
  • Real User Experience Insight [Product ID 9572/COLL]
  • SAM-QFS [Product ID 10021]
  • Scapp [Product ID 9851]
  • Siebel CRM [Product ID 2295]
  • SMS [Product ID 9852]
  • Solaris 11.1 and before [Product ID 10006]
  • SPARC - OPL Service Processor (XCP) [Product ID 9845]
  • Sun Blade 6000 Ethernet Switched NEM 24P 10GE [Product ID 9889/OPUS-TOR]
  • Sun Crypto Accelerator 6000 [Product ID 9894]
  • Sun Network 10GE Switch 72p [Product ID 9889/OPUSC10NEM]
  • Sun Ray Operating Software 11.x [Product ID 9211]
  • Sun Ray Software 5.x [Product ID 8242]
  • Sun System Firmware [Product ID 9846]
  • Tape Library SL500, SL3000, SL8500 [Product ID 10101, 10100, 10102]
  • Tekelec HLR Router 3.0.0 , 3.1.0 [Product ID None yet]
  • Tekelec Platform Management and Configuration 5.0.0, 5.5.0 [Product ID None yet]
  • Webgate 10g and 11g
  • Webtier "the old" 1.0.2.2 webtier (E-Business Suite uses) [Product ID 1042/MODSSL]
Not certain, Products which are still under investigation :

  • ATG Products
  • eGate Integrator 5.0.5 SRE
  • Enterprise Manager Base Platform [Product ID 1370]
  • Enterprise Manager Explorer
  • Entperprise Manager TDS
  • FLEXCUBE Connect [Product ID 9051]
  • Java CAPS 6.2 [Product ID 8528]
  • MySQL Connector/C++ [Product ID 8576/CONCPLS]
  • Nimbula Director [Product ID 10773]
  • OnTrack Release
  • Oracle GoldenGate
  • Oracle Service Bus [Product ID 5308]
  • Oracle SOA Suite [Product ID 1162]
  • Oracle Wallet Manager [Product ID 338,991/WMT]
  • SuperCluster [Product ID 10011]
Affected products which have fixes:

Patch Availability Matrix
Affected Products
Patch Availability
MySQL Enterprise Monitor [Product ID 8480]Contact Support
MySQL Enterprise Server version 5.6 [Product ID 8476]Customers can download version 5.6.18 to address CVE-2014-0160. The download instructions can be found at
https://support.oracle.com/epmos/faces/DocumentDisplay?id=1300654.1
Oracle Communications Session Monitor Suite 3.3.40, 3.3.50 [Product ID 10761]Patchset available. Contact Support. 3.3.40.2.1 (available as of April 11th, 2014). 3.3.50.0.0 planned.
Oracle Linux 6https://linux.oracle.com/cve/CVE-2014-0160.html
https://linux.oracle.com/errata/ELSA-2014-0376.html
Note that Oracle Linux 6 ships a patched version of OpenSSL 1.0.1e on the Unbreakable Linux Network and public-yum.oracle.com. Customers are strongly encouraged to upgrade OpenSSL on Oracle Linux 6 to the latest available release.
Oracle Mobile Security Suite [Product ID 10913]1. Login to support.oracle.com
2. Select "Patches & Updates"
3. Search for the appropriate patch by Bug Number
  - For the patch on top of OMSS v2.5.x, search for Bug Number 18545175
  - For the patch on top of OMSS v3.0.x, search for Bug Number 18545252
4. Download the appropriate patch(es)
5. Follow the instructions in the readme.txt contained in the patch zip file
Solaris 11.2 (Selected Customers Only) [Product ID 10026]Contact Support


Affected products which have no fixes:

  • BlueKai
  • Java ME - JSRs and Optional Packages
  • Java ME - Mobile and Wireless
  • MySQL Connector/C [Product ID 8576/CONC]
  • MySQL Connector/ODBC [Product ID 8576/CONODBC]
  • MySQL Workbench [Product ID 4627]
  • Oracle Communications Internet Name and Address Management [Product ID 2262]
  • Oracle Communications Application Session Controller 3.7.0m2p0 [Product ID 10769]
  • Oracle Communications Interactive Session Recorder 4.x, 5.0, 5.1 [Product ID 10765]
  • Oracle Communications Network Charging and Control [Product ID 4623]
  • Oracle Communications Policy Management 11.1 [Product ID None yet]
  • Oracle Communications Session Delivery Management Suite NNC 7.3 [Product ID None yet]
  • Oracle Communications WebRTC Session Controller/Director WSC 7.0.1 [Product ID 10811]
  • Primavera P6 Professional Project Management [Product ID 5580]

Products which dont include Open SSL:

  • Auto Service Request [Product ID 9042]
  • E-Business Suite R12
  • FLEXCUBE Messaging Hub [Product ID 9102]
  • FLEXCUBE Remit [Product ID 9097]
  • Hyperion EPM
  • Java ME - Bluray and TV [Product ID 9319]
  • Java ME - Embedded [Product ID 9326]
  • Java ME - Javacard [Product ID 9328]
  • Java SE [Product ID 856]
  • JavaVM
  • MySQL Cluster [Product ID 8479]
  • MySQL Cluster Manager [Product ID 8479/CLSTMGR]
  • MySQL Community Server version 5.6 [Product ID 6850]
  • MySQL Connector/Java [Product ID 8576/CONJ]
  • MySQL Connector/NET [Product ID 8576/CONNET]
  • MySQL Connector/PHP (mysqlnd) [Product ID 8576/CONMYND]
  • MySQL Connector/Python [Product ID 8576/CONPYTHN]
  • MySQL Server (all licenses, versions 5.5 and earlier) [Product ID 8478]
  • MySQL Utilities [Product ID 4627/WBUTILS]
  • OFSS FLEXCUBE Electronic Bill Presentment and Payment [Product ID 9495]
  • Oracle Access Manager 11g Server
  • Oracle Access Portal [Product ID 10878]
  • Oracle Adaptive Access Manager 11g Server [Product ID 4419]
  • Oracle Agile Product Lifecycle Management [Product ID 4461]
  • Oracle Application Express (formerly Oracle HTML DB) [Product ID 1348]
  • Oracle Application REST Data Services (formerly Oracle APEX Listener) [Product ID 9456]
  • Oracle Autovue
  • Oracle Banking Platform [Product ID 9178]
  • Oracle Business Intelligence Discoverer [Product ID 964]
  • Oracle CODASYL DBMS [Product ID 624]
  • Oracle Communications Configuration Management 7.2 [Product ID 2268]
  • Oracle Communications Eagle STP 44.0.x, 45.0.x
  • Oracle Communications Elastic Charging Engine [Product ID 9742]
  • Oracle Communications Offline Mediation Controller [Product ID 2269]
  • Oracle Communications Pricing Design Center [Product ID 9437]
  • Oracle Complex Maintenance, Repair and Overhaul [Product ID 1184]
  • Oracle Database [Product ID 5]
  • Oracle Depot Repair [Product ID 516]
  • Oracle Directory Server Enterprise Edition [Product ID 8512]
  • Oracle Documaker [Product ID 5477]
  • Oracle Enterprise Limits and Collateral
  • Oracle Enterprise Single Sign-On Suite [Product ID 2074]
  • Oracle Financial Service Lending and Leasing [Product ID 10484]
  • Oracle FLEXCUBE Core Banking [Product ID 9101]
  • Oracle FLEXCUBE Direct Banking [Product ID 9111]
  • Oracle FLEXCUBE Investment Services [Product ID 9099]
  • Oracle FLEXCUBE Private Banking [Product ID 9110]
  • Oracle FLEXCUBE Universal Banking [Product ID 9052]
  • Oracle GlassFish Communications Server 2.x [Product ID 8513]
  • Oracle Health Sciences InForm Adapter [Product ID 9637]
  • Oracle Health Sciences InForm and Oracle Siebel Clinical Integration Pack for Subject and Status Information [Product ID 9601]
  • Oracle Health Sciences InForm CRF Submit [Product ID 9641]
  • Oracle Health Sciences InForm Publisher [Product ID 9638]
  • Oracle HTTP Server [Product ID 1042]
  • Oracle Identity Analytics [Product ID 8522]
  • Oracle Identity Federation 11g [Product ID 1741]
  • Oracle Identity Manager [Product ID 1980]
  • Oracle Internet Directory [Product ID 355]
  • Oracle iPlanet Web Proxy Server 4.0 [Product ID 8542]
  • Oracle iPlanet Web Server 7.0 [Product ID 8543]
  • Oracle Knowledge [Product ID 9571]
  • Oracle Mobile and Social [Product ID 9146]
  • Oracle Pedigree and Serialization Manager [Product ID 4674]
  • Oracle Portal [Product ID 96]
  • Oracle Rdb Server on OpenVMS [Product ID 623]
  • Oracle Security Token Service 11g [Product ID 5744]
  • Oracle Sun OpenSSO 8.x Server [Product ID 8520]
  • Oracle Trace File Analyzer
  • Oracle Traffic Director [Product ID 9276]
  • Oracle Transportation Management 6.3 [Product ID 1991]
  • Oracle Unified Directory [Product ID 9118]
  • Oracle Virtual Desktop Client 3.x [Product ID 8541]
  • Oracle Waveset [Product ID 8518]
  • Oracle Web Cache [Product ID 1059]
  • Oracle WebLogic Server [Product ID 5242]
  • Oracle WebLogic Web Server Plug-In 1.1+, 11g, 12c [Product ID 5242/PLUGIN_NZ]
  • Retail Integration Bus [Product ID 1807]
  • StorageTek Automated Cartridge System Library Software [Product ID 10088]
  • Sun GlassFish Enterprise Server 2.x
  • Sun Java System Application Server 7.x, 8.x [Product ID 6802]
  • Sun Java System Web Proxy Server 3.6+, 4.0+
  • Sun Java System Web Server 7.0 [Product ID 7276]
  • Sun ONE Web Server 6.1 [Product ID 8543]
  • Sun Storage Common Array Manager (CAM) [Product ID 10024]
Oracle Cloud Services, which are not affected:

  • Oracle Public Cloud
    • RightNow
    • Big Machines
    • Eloqua
    • Responsys
  • Oracle Managed Cloud Services
    • All Services
  • Oracle Cloud for Industry
    • Argus Safety, Insight, Analytics
    • Central Coding
    • Central Designer
    • ClearTrial
    • CTMS
    • Empirica Suite
    • Healthcare Analytic Suite
    • InForm
    • IRT
    • LabPas
    • Outcome Logix
    • Insurance Data Exchange
    • Enterprise Track (Instantis)
    • Skire Unifier
    • Primavera P6
    • Oracle Utilities Cloud Analytics (previously DataRaker)
    • Billing and Revenue Management
    • Oracle Financial Services Lending and Leasing