Showing posts with label RMS. Show all posts
Showing posts with label RMS. Show all posts

Friday, October 19, 2012

[OpsMgr 2007R2][OpsMgr 2012] Troubleshooting management server and gateway performance

 

For OpsMgr 2007 and R2 - Root management server (RMS)

Configuration update bursts are caused by management pack imports and by discovery data. When system performance is slow, the most likely bottlenecks are, first, the CPU and, second, the OpsMgr installation disk I/O.

The RMS is responsible for generating and sending configuration files to all affected Health Services.

For Workflow reloading (which is caused by new configuration on RMS), the most likely bottlenecks are the same: the CPU first, and OpsMgr installation disk I/O second. The RMS is responsible for reading the configuration file, for loading and initializing all workflows that run on it, and for updating the RMS HealthService store when the configuration file is updated on the RMS.

For local workflow activity bursts (which is when agents change their availability), the most likely bottleneck is the CPU. If you find that the CPU is not working at maximum capacity, the next most likely bottleneck is the hard disk. The RMS is responsible for monitoring the availability of all agents that are using RMS local workflows. The RMS also hosts distributed dependency monitors that use the disk.

Management server

During a configuration update burst (that is caused by MP import and discovery), the typical bottlenecks are, first, the CPU and, second, the OpsMgr installation disk I/O. The management server is responsible of forwarding configuration files from the RMS to the target agents.

For Operational data collection, bottlenecks are typically caused by the CPU. The disk I/O may also be at maximum capacity, but that is not as likely. The management server is responsible for decompressing and decrypting incoming operational data, and inserting it into the Operational Database. It also sends acknowledgements (ACKs) back to the agents or gateways after it receives operational data, and uses disk queuing to temporarily store these outgoing ACKs. Lastly, the management server will also forward monitor state changes (by using a disk queue) to the RMS for distributed dependency monitors.

Gateway

The gateway is both CPU-bound and I/O-bound. When the gateway is relaying a large amount of data, both the CPU and I/O operations may show high usage. Most of the CPU usage is caused by the decompression, compression, encryption, and decryption of the incoming data, and also by the transfer of that data. All data that is received by the gateway and from the agents is stored in a persistent queue on disk, to be read and forwarded to the management server by the gateway Health service. This can cause heavy disk usage. This usage can be significant when the gateway is taken temporarily offline and must then handle accumulated agent data that the agents generated and tried to send when the GW was still offline.

To troubleshoot the issue in this situation, collect the following information for each affected management server or gateway:
  • Exact Windows version, edition, and build number (for example, Windows Server 2003 Enterprise x64 SP2)
  • Number of processors
  • Amount of RAM
  • Drive that contains the Health Service State folder
  • Whether the antivirus software is configured to exclude the Health Service store

    Note For more information, click the following article number to view the article in the Microsoft Knowledge Base: 975931 (http://support.microsoft.com/kb/975931/ )

  • Recommendations for antivirus exclusions that relate to Operations Manager
  • RAID level (0, 1, 5, 0+1 or 1+0) for the drive that is used by the Health Service State
  • Number of disks used for the RAID
  • Whether battery-backed write cache is enabled on the array controller

This posting is provided "AS IS" with no warranties.

Monday, September 10, 2012

[OpsMgr 2007][OpsMgr 2012] Usefull SQL Query / Report to retrieve what is targeting the RMS

has just published a SQL query and a report that will be very helpfull for the one that is preparing a migration to SCOM 2012.

Retrieving rules, monitors, tasks and discoveries that target the RMS will be usefull to decide what to do know :
  • if you have to modify your own management packs to target MS instead of RMS,
  • what provided MP is targeting the RMS (Microsoft Exchange Server 2010 Management Pack, Citrix EdgeSight Management Pack...) and also help you to decide if you have to install the RMS emulator in your new SCOM 2012 environment.


Please find his SQL query that must be launched on the Operations Manager DB.

DECLARE @ManagedType varchar(200)
SET @ManagedType = 'Root Management Server'
SELECT     ManagedTypeView.DisplayName, ManagementPackView.DisplayName AS 'Management Pack', 'Monitors' AS 'Type', ManagedEntityGenericView.DisplayName AS 'Target',
                      MonitorView.DisplayName AS 'Name', MonitorView.Category, MonitorView.Description,
                      CASE MonitorView.Enabled
                      WHEN '0' THEN 'Disabled'
                      WHEN '2' THEN 'Enabled'
                      WHEN '3' THEN 'Enabled'
                      WHEN '4' THEN 'Enabled'
                      End As 'Enabled', CONVERT(VARCHAR(20),
                      MonitorView.TimeAdded, 102) AS 'TimeAdded'
FROM         ManagementPackView INNER JOIN
                      MonitorView WITH(NOLOCK) ON ManagementPackView.Id = MonitorView.ManagementPackId INNER JOIN
                      ManagedEntityGenericView WITH(NOLOCK) ON MonitorView.TargetMonitoringClassId = ManagedEntityGenericView.MonitoringClassId INNER JOIN
                      TypedManagedEntity  WITH(NOLOCK) ON ManagedEntityGenericView.TypedManagedEntityId = TypedManagedEntity.TypedManagedEntityId INNER JOIN
                      ManagedTypeView  WITH(NOLOCK) ON TypedManagedEntity.ManagedTypeId = ManagedTypeView.Id
WHERE     (ManagedTypeView.DisplayName like '%' + @ManagedType + '%')
UNION
SELECT      ManagedTypeView_1.DisplayName, ManagementPackView_1.DisplayName AS 'Management Pack', 'Rules' AS 'Type', ManagedEntityGenericView_1.DisplayName AS 'Target',
                      RuleView.DisplayName AS 'Name', RuleView.Category, RuleView.Description, CASE RuleView.Enabled
                      WHEN '0' THEN 'Disabled'
                      WHEN '2' THEN 'Enabled'
                      WHEN '3' THEN 'Enabled'
                      WHEN '4' THEN 'Enabled'
                      End As 'Enabled', CONVERT(VARCHAR(20), RuleView.TimeAdded, 102)
                      AS 'TimeAdded'
FROM         ManagementPackView AS ManagementPackView_1 INNER JOIN
                      RuleView WITH(NOLOCK) ON ManagementPackView_1.Id = RuleView.ManagementPackId INNER JOIN
                      ManagedEntityGenericView AS ManagedEntityGenericView_1 WITH(NOLOCK) ON
                      RuleView.TargetMonitoringClassId = ManagedEntityGenericView_1.MonitoringClassId INNER JOIN
                      TypedManagedEntity AS TypedManagedEntity_1 WITH(NOLOCK) ON ManagedEntityGenericView_1.TypedManagedEntityId = TypedManagedEntity_1.TypedManagedEntityId INNER JOIN
                      ManagedTypeView As ManagedTypeView_1 WITH(NOLOCK) ON TypedManagedEntity_1.ManagedTypeId = ManagedTypeView_1.Id
WHERE     (ManagedTypeView_1.DisplayName like '%' + @ManagedType + '%')
UNION
SELECT      ManagedTypeView_2.DisplayName, ManagementPackView_2.DisplayName AS 'Management Pack', 'Discoveries' AS 'Type', ManagedEntityGenericView_2.DisplayName AS 'Target',
                      DiscoveryView.DisplayName AS 'Name', case DiscoveryView.Category when 12 THEN 'Discovery' END, DiscoveryView.Description, CASE DiscoveryView.Enabled
                      WHEN '0' THEN 'Disabled'
                      WHEN '2' THEN 'Enabled'
                      WHEN '3' THEN 'Enabled'
                      WHEN '4' THEN 'Enabled'
                      End As 'Enabled', CONVERT(VARCHAR(20), DiscoveryView.TimeAdded, 102)
                      AS 'TimeAdded'
FROM         ManagementPackView AS ManagementPackView_2 INNER JOIN
                      DiscoveryView WITH(NOLOCK) ON ManagementPackView_2.Id = DiscoveryView.ManagementPackId INNER JOIN
                      ManagedEntityGenericView AS ManagedEntityGenericView_2 WITH(NOLOCK) ON
                      DiscoveryView.TargetMonitoringClassId = ManagedEntityGenericView_2.MonitoringClassId INNER JOIN
                      TypedManagedEntity AS TypedManagedEntity_2 WITH(NOLOCK) ON ManagedEntityGenericView_2.TypedManagedEntityId = TypedManagedEntity_2.TypedManagedEntityId INNER JOIN
                      ManagedTypeView As ManagedTypeView_2 WITH(NOLOCK) ON TypedManagedEntity_2.ManagedTypeId = ManagedTypeView_2.Id
WHERE     (ManagedTypeView_2.DisplayName like '%' + @ManagedType + '%')
UNION
SELECT      ManagedTypeView_3.DisplayName, ManagementPackView_3.DisplayName AS 'Management Pack', 'Tasks' AS 'Type', ManagedEntityGenericView_3.DisplayName AS 'Target',
                      RecoveryView.DisplayName AS 'Name', case RecoveryView.Category  when 17 THEN 'Recovery' END, RecoveryView.Description, CASE RecoveryView.Enabled
                      WHEN '0' THEN 'Disabled'
                      WHEN '2' THEN 'Enabled'
                      WHEN '3' THEN 'Enabled'
                      WHEN '4' THEN 'Enabled'
                      End As 'Enabled', CONVERT(VARCHAR(20), RecoveryView.TimeAdded, 102)
                      AS 'TimeAdded'
FROM         ManagementPackView AS ManagementPackView_3 INNER JOIN
                      RecoveryView WITH(NOLOCK) ON ManagementPackView_3.Id = RecoveryView.ManagementPackId INNER JOIN
                      ManagedEntityGenericView AS ManagedEntityGenericView_3 WITH(NOLOCK) ON
                      RecoveryView.TargetMonitoringClassId = ManagedEntityGenericView_3.MonitoringClassId INNER JOIN
                      TypedManagedEntity AS TypedManagedEntity_3 WITH(NOLOCK) ON ManagedEntityGenericView_3.TypedManagedEntityId = TypedManagedEntity_3.TypedManagedEntityId INNER JOIN
                      ManagedTypeView As ManagedTypeView_3 WITH(NOLOCK) ON TypedManagedEntity_3.ManagedTypeId = ManagedTypeView_3.Id
WHERE     (ManagedTypeView_3.DisplayName like '%' + @ManagedType + '%')
ORDER BY 'Management Pack'







This posting is provided "AS IS" with no warranties.

Friday, April 27, 2012

[OpsMgr 2007] How to troubleshoot Event ID 2115 in Operations Manager

Article ID: 2681388 - Last Review: April 17, 2012 - Revision: 1.1
 
Microsoft has published KB2681388 to expose how to troubleshoot Event ID 2115 on your RMS

Symptomes :
In Operations Manager, one of the performance concerns surrounds Operations Manager Database and Data Warehouse insertion times. The following is a description to help identify and troubleshoot problems concerning Database and Data Warehouse data insertion.

Examine the Operations Manager Event log for the presence of Event ID 2115 events. These events typically indicate that performance issues exist on the Management Server or the Microsoft SQL Server that is hosting the OperationsManager or OperationsManager Data Warehouse databases. Database and Data Warehouse write action workflows run on the Management Servers and these workflows first retain the data received from the Agents and Gateway Servers in an internal buffer. They then gather this data from the internal buffer and insert it into the Database and Data Warehouse. When the first data insertion has completed, the workflows will then create another batch.

The size of each batch of data depends on how much data is available in the buffer when the batch is created, however there is a maximum limit on the size of the data batch of up to 5000 data items.  If the data item incoming rate increases, or the data item insertion throughput to the Operation Manager and Data Warehouse databases throughput is reduced, the buffer will then accumulate more data and the batch size will grow larger.  There are several write action workflows that run on a Management Server.  These workflows handle data insertion to the Operations Manager and Data Warehouse databases for different data types.  For example:
  • Microsoft.SystemCenter.DataWarehouse.CollectEntityHealthStateChange
  • Microsoft.SystemCenter.DataWarehouse.CollectPerformanceData
  • Microsoft.SystemCenter.DataWarehouse.CollectEventData
  • Microsoft.SystemCenter.CollectAlerts
  • Microsoft.SystemCenter.CollectEntityState
  • Microsoft.SystemCenter.CollectPublishedEntityState
  • Microsoft.SystemCenter.CollectDiscoveryData
  • Microsoft.SystemCenter.CollectSignatureData
  • Microsoft.SystemCenter.CollectEventData

When a Database or Data Warehouse write action workflow on a Management Server experiences slow data batch insertion, for example times in excess of 60 seconds, it will begin logging Event ID 2115 to the Operations Manager Event log. This event is logged every one minute until the data batch is inserted into the Database or Data Warehouse, or the data is dropped by the write action workflow module. As a result, Event ID 2115 will be logged due to the latency inserting data into the Database or Data Warehouse. Below is an example Event logged due to data dropped by the write action workflow module: 

Event Type: Error
Event Source: HealthService
Event Category: None
Event ID: 4506
Computer: <RMS NAME>
Description:
Data was dropped due to too much outstanding data in rule "Microsoft.SystemCenter.OperationalDataReporting.SubmitOperationalDataFailed.Alert" running for instance <RMS NAME> with id:"{F56EB161-4ABE-5BC7-610F-4365524F294E}" in management group <MANAGEMENT GROUP NAME>.


Event ID 2115 contains 2 significant pieces of information.  First, the name of the Workflow that is experiencing the problem and second, the elapsed time since the workflow began inserting the last batch of data. 

For example:

Log Name: Operations Manager
Source:        HealthService
Event ID:      2115
Level:         Warning
Computer:      <RMS NAME>
Description:
A Bind Data Source in Management Group <MANGEMENT GROUP NAME> has posted items to the workflow, but has not received a response in 300 seconds.  This indicates a performance or functional problem with the workflow.
 Workflow Id : Microsoft.SystemCenter.CollectPublishedEntityState
 Instance    : <RMS NAME>
Instance Id : {88676CDF-E284-7838-AC70-E898DA1720CB}


This particular Event ID 2115 message indicates that the workflow Microsoft.SystemCenter.CollectPublishedEntityState, which writes Entity State data to the Operations Manager database, is trying to insert a batch of Entity State data and it started 300 seconds ago.  In this example the insertion of the Entity State data has not yet finished.  Normally inserting a batch of data should complete within 60 seconds.  If the Workflow Id contains Data Warehouse then the problem concerns the Operations Manager Data Warehouse.  Otherwise, the problem would concern inserting data into the Operations Manager database.

Cause :

As the description of Event ID 2115 states, this may indicate a database performance problem or too much data incoming from the agents. Event ID 2115 simply indicates there is a backlog inserting data into the Database; Operations Manager or Operations Manager Data Warehouse. These Events can originate from a number of possible causes. For example, a large amount of Discovery data, a Database connectivity issue or full database condition, potential disk or network constraints.

In Operations Manager, Discovery data insertion is a relatively expensive process. We define a burst of data as a short period of time where a significant amount of data is received by the Management Server. These bursts of data can cause Event ID 2115 since the data insertion should occur infrequently. If Event ID 2115 consistently appears for Discovery data collection, this can indicate either a Database or Data Warehouse insertion problem or Discovery rules in a Management Pack collecting too much discovery data.

Operations Manager configuration updates caused by Instance Space changes or Management Pack imports have a direct effect on CPU utilization on the Database Server and this can impact Database insertion times. Following a Management Pack import or a large instance space change, it is expected to see Event ID 2115 messages. For more information on this topic please see the following:

2603913 - How to detect and troubleshoot frequent configuration changes in Operations Manager (http://support.microsoft.com/kb/2603913 (http://support.microsoft.com/kb/2603913)

If the Operations Manager or Operations Manager Data Warehouse databases are out of space or offline, it is expected that the Management Server will continue to log Event ID 2115 messages to the Operations Manager Event log and the pending time will grow higher.

If the write action workflows cannot connect to the Operations Manager or Operations Manager Data Warehouse databases, or they are using invalid credentials to establish their connection, the data insertion will be blocked and Event ID 2115 messages will be logged accordingly until this situation is resolved.

In Operations Manager, expensive User Interface queries can impact resource utilization on the Database which can lead to latency in Database insertion times. When a user is performing an expensive User Interface operation it is possible to see Event ID 2115 messages logged.


Event ID 2115 messages can also indicate a performance problem if the Operations Manager Database and Data Warehouse databases are not properly configured. Performance problems on the database servers can lead to Event ID 2115 messages. Some possible causes include the following:
  • The SQL Log or TempDB database is too small or out of space.
  • The Network link from the Operations Manager and Data Warehouse database to the Management Server is bandwidth constrained or the latency is large. In this scenario we recommend to Management Server to be on the same LAN as the Operations Manager and Data Warehouse server.
  • The data disk hosting the Database, logs or TempDB used by the Operations Manager and Data Warehouse databases is slow or experiencing a function problem. In this scenario we recommend leveraging RAID 10 and we also recommend enabling battery backed Write Cache on the Array Controller.
  • The Operations Manager Database or Data Warehouse server does not have sufficient memory or CPU resources.
  • The SQL Server instance hosting the Operations Manager Database or Data Warehouse is offline.

It is recommend that the Management Server reside on the same LAN as the Operations Manager and Data Warehouse database server.

Event ID 2115 messages can also occur if the disk subsystem hosting the Database, logs or TempDB used by the Operations Manager and Data Warehouse databases is slow or experiencing a function problem. In this scenario we recommend leveraging RAID 10 and we also recommend enabling battery backed Write Cache on the Array Controller. 

Resolution :
Microsoft propose 6 scenarios to solve the issue.



This posting is provided "AS IS" with no warranties.