Showing posts with label What's New. Show all posts
Showing posts with label What's New. Show all posts

Monday, June 26, 2017

Part 5 - Performance Troubleshooting Dashboards in vRealize Operations 6.6.

Hope you are enjoying the What's New with vROps 6.6 Series. I am having a great time writing this, since my experience as a user of vROps has completely turned around with this release. In this post, we will continue talking about the rich & use case driven out of the box content available in the form of dashboards.

The Getting Started page in the product acts an anchor for showcasing all the use cases. The focus of this post would be on Performance Troubleshooting.

Here is how Performance Troubleshooting shows up on the Getting Started Page:


The Performance Troubleshooting category caters to the administrators responsible for managing the performance & availability of the virtual machines running in the virtual infrastructure. This category runs your through a guided workflow to answer questions which will help you with the troubleshooting process. The dashboards in this category identify and isolate problems that may impact your applications. They provide a line of sight into the full stack to isolate and identify the root cause quickly.


Key questions these dashboards help you answer are :

Is application performance impacted due to virtual infrastructure?

Are noisy neighbors impacting multiple virtual machines and corresponding applications?

Are there active alerts which require action?

Any known issues impacting the performance & availability of a vSAN cluster?


Let us look at each of these dashboard and I will provide a summary of what these dashboards can do for you along with a quick view of the dashboard:



Troubleshoot a VM

The Troubleshoot a VM dashboard helps a VI Administrator to deal with day to day troubleshooting of issues in a virtual infrastructure. While most of the IT issues in an organization are reported at the application layer, this dashboard provides a guided workflow which can help investigate an ongoing or a suspected issue with virtual machines supporting the impacted applications.

You can easily search for a virtual machine by its name or can sort the list of VMs with active alerts on them to start your troubleshooting process. As soon as you select a VM, you can view its key properties to ensure they are configured as per your virtual infrastructure design. Any deviation from standards could cause potential issues. You can view any known alerts, the workload trend of the VM over the past week and if any of the resources serving the virtual machine has any ongoing issue.

The next step in the troubleshooting process allows you to eliminate the major symptoms which might impact the performance or availability of a virtual machine. You can drill down further into the key metrics to find out if the VMs utilization patterns are abnormal or it is contending for basic resources such as CPU, Memory or Disk.




Troubleshoot a Cluster

The Troubleshoot a Cluster Dashboard provides you a guided workflow to identify issues and isolate them easily. You can either start with a cluster which happens to have an issue by using the search option or you can simply sort your clusters with the number of active alerts on them.

On selecting a specific cluster you want to work with, you can see a quick summary of number of hosts participating in that cluster and the VMs being served by them. The dashboard provides you the current and past utilization trends of how hard your cluster is working and what are the known problems on the cluster in form of alerts.

You can easily view the hierarchy of objects related to the cluster and review their status to identify if they are impacted due to the current health of the cluster. You can quickly identify any contention issues by looking at the max and avg. contention faced by the virtual machines on the selected cluster. The dashboard allows you to drill down to specific virtual machines which might be a victim of resource contention and take your next steps in the troubleshooting process to cater to those victims and avoid issues proactively. 



Troubleshoot a Datastore

Troubleshoot a Datastore dashboard helps provide a guided workload to an administrator in order to quickly identify storage issues and act on them. Based on your troubleshooting style you can either start with a Datastore which might be in trouble due to high latency and is showing red on the heatmap or you can search for a Datastore which you have in mind. You can also sort all the datastores with active alerts and start working your way with a Datastore with known problems.

On selecting a datastore you see its current capacity and utilization along with a count of VMs served by that Datastore. The metric charts helps you to view historical trends of key storage metrics such as latency, outstanding IOs and throughput.

The dashboard also lists the virtual machines served by the selected datastore and help you analyze the utilization and performance trends of those virtual machines. If the virtual machines are suffering, the VI administrator can migrate these virtual machines over to other datastores to evenly spread out IO load.



Troubleshoot a Host

Since ESXi servers are the main source of providing resources to a virtual machine, they become extremely critical when it comes to performance and availability. With Troubleshoot a host dashboard, you can either search for specific Host which you have in mind or sort the hosts with active alerts to start your investigation.

As soon as you select a host, you can see the key properties of each of the host to ensure thy are configured as per your virtual infrastructure design. Any deviation from standards could cause potential issues. You can answer some key questions around current and past utilization, workload trends over the last week and if virtual machines served by the host are healthy.

Hardware faults with the hosts can be easily surfaced on this dashboard since it lists all the critical events which might affect the availability of the hosts. If you find a host which is running hot, the next logical step in the troubleshooting process would be to find out the villain virtual machines which might be consuming resources from that host. You can find a list of top 10 virtual machines which are demanding CPU and Memory Resources from the identified host.




Troubleshoot vSAN

The Troubleshoot vSAN dashboard is designed to help a vSAN administrator step through a guided workflow to investigate potential issues with each layer of vSAN. The dashboard allows you to start with looking at key properties of your vSAN cluster along with the active alerts on any of the cluster components such as hosts, disk groups or the vSAN datastores.

Once you select a cluster, you can list all the known problems with all the objects which are associated to that cluster. This includes, clusters, datastores, disk groups, physical disks and most importantly the virtual machines which are being served by the selected vSAN cluster.

The dashboard then drills down into the key utilization and performance metrics and shows you a trend of how the cluster has been used and has performed over the past 24 hours. You can easily go back in time if you are dealing with historical issues. While most of the problems would be surfaced up at the cluster level, a drill down analysis can be done at the host, disk group or down to the physical disk level.

Heatmaps within the dashboard help you answer questions around write buffer usage, cache hit ratio, host configurations and physical issues with capacity and cache disks such as drive wearout, drive temperature and read-write errors.





 Troubleshoot with Logs

The Troubleshoot with Logs dashboard can be used when you want to investigate an ongoing issue within your virtual infrastructure using the logs. This dashboard helps you to look at predefined views created within Log Insight to answer common questions which can be surfaced through pre-defined queries within Log insight.

With this dashboard, you can correlate metrics and queries within vRealize Operations Manager on a single pane of glass to troubleshoot issues across applications and infrastructure.





In case you are like me, and don't like to READ. You can see the dashboards in action in this video playlist:

See all dashboards in action here.



More to come.. Stay Tuned!!




Tuesday, March 21, 2017

Did You Know #6 - Using custom metrics groups in vROps for troubleshooting


Welcome back to the Did You Know Series on vRealize operations Manager. As I mentioned in the first part of this series, the goal here is to unearth the Best Kept Secrets of vRealize Operations Manager. 

This could be across features, functionalities, use cases, integrations, APIs or any tips or tricks which can help make day to day operations of SDDC easier and fun with vROps!

Today I will talk about how you can make troubleshooting an easier process with vROps. Troubleshooting as we all know is not a skill, it is a methodology. It would NOT be incorrect to say that each one of us has a different troubleshooting style. With vRealize Operations you can troubleshoot issues the way you like them. Some prefer OOTB dashboards, some like to create there own personalized views and some prefer to jump into what me and Iwan call God Mode aka the All Metric view in the product.

The All Metrics view of the product can easily become complex as it shows you all the metrics which are associated to an Object Type. If you look at vRealize Operations Manager 6.4, the product gave you some OOTB custom metric groups which can be used to list all common metrics around CPU, Memory and Disk. These were the OOTB options and might not fit all the needs. If you are on vROps 6.4, click on a VM object and click on All Metrics and you will see this:


You can see that apart from all metrics and all properties for the virtual machine, you can see 5 custom categories which list specific metrics which can be used for troubleshooting. As soon as I double click on CPU, I will see all the key metrics pertaining to the CPU Metric group in one shot on the right pane:



While this is a cool feature and make troubleshooting really simple, there was one use case which could not be solved here. If as an admin I wanted to create my own metric group where I want to focus on key metrics of my own choice, I was unable to create a metric group for same. While this was not possible with vROps 6.4, with the arrival of vROps 6.5, this feature is now available and now you can create your own custom metric groups with the metrics you like to use for troubleshooting.

I will show you how to create one custom group at virtual machine level:

1- Click on any virtual machine in your environment.

2- Click on All Metrics.

3- Click on the blue wheel and click on the Add Group option.















4- Provide a name. I will call it "VM KPIs"









5- Now you can drag and drop any metric from the all metrics group to this new created metric group:



Here are a few KPIs which I added in my vROps as they help me troubleshoot in God Mode with a single click.....

VM KPIs

CPU | Demand %
CPU | Usage %
CPU | CPU Contention %
CPU | CO-Stop %
CPU | Ready %

Memory | Usage %
Memory | Contention %
Memory | Balloon %
Memory | Swap In (KB)
Memory | Compressed (KB)

Virtual Disk | Aggregate of all instances | Commands Per Second
Virtual Disk | Aggregate of all instances | Total Latency
Disk Space | Snapshot | Virtual Machine Used (GB)
Guest File System Stats | Total Guest File System Free (GB)

Network I/O|Aggregate of all instances| Packets Dropped %
















Host System KPIs

CPU|CPU Contention (%)
CPU|Demand (%)
Memory|Contention (%)
Memory|Total Capacity (KB)
Memory|Consumed (KB)
Memory| Usage (%)
Network I/O|Aggregate of all instances|Packets Dropped (%)












Cluster KPIs

CPU| CPU Contention (%)
CPU|Demand (%)
CPU|Max VM CPU Contention (%)
Memory|Balloon (KB)
Memory|Contention (%)
Memory|Max VM memory Contention (%)
Memory|Usage (%)



 Go configure your vROps with the metrics you like and make troubleshooting an easy and fun process...

And yeah. Keep sharing!!




Monday, October 26, 2015

Virtual Machine Capacity Profiles in vRealize Operations 6.1

Let me start this article with this statement, "vRealize Operations Manager 6.1 is probably the best UI for an enterprise Performance, Capacity or Monitoring tool I have ever seen". I must congratulate the VMware R&D folks and the Product Management to transform vCenter Operations to vRealize Operations in a beautiful way. Yes I work for VMware and this might sound bias, but I would encourage you to use the product once and I am sure you would agree with me about this transformation for good.

While the UI has become amazingly intuitive, the great news is that simplification of the product from a feature-functionality standpoint is definitely another area where I am liking the product more than what I use to. One of such a feature is the Capacity Remaining Breakdown.

We all know that vROps has the capability of trend analysis and forecasting for capacity on the basis of policies which are pre-defined by IT for an organization or for an environment. One of the outcomes of the policies is to derive the CAPACITY REMAINING. This derived metric uses either the Demand or Allocation model for capacity planning and determines the number of VMs Remaining (in other words, the number of VMs which can be deployed on a container before it is declared full). A container here is a logical abstract of resources and are denoted by an ESXi host, a resource pool, a cluster, a virtual datacenter, a vCenter Server or it could also be the entire Universe.

If you remember the One Click cluster capacity dashboard which I created for vCOps 5.x here and then re-created for vROps 6.x here, you would notice that I plot the VM Remaining value which denotes the number of VMs which are left in a given cluster. Here are the screenshots where I am pointing at both the dashboards with a little explanation:


Dashboard with vCOps 5.x

If you notice the highlighted fields in RED, you can clearly see that, I am showcasing the number of VMs left in the cluster & I am showcasing the number of VMs left in the cluster resource wise, such as CPU, Memory, Disk & Network. In this case, the number of VMs remaining is calculated using the following formula:

VM Remaining = (Total Resources Average VM Profile Size) - VMs Deployed

Remember the Average VM Profile Size here is automatically determined by vROps and hence the VMs Remaining would be of that Profile Size. This size can bee seen easily under the analysis tab of vROps/vCOps.



Dashboard with vROps 6.x

Now lets's look at the Dashboard which I created for version 6.x (to be precise any version of vROps 6.x except 6.1).

Here you can notice that apart from the VM Remaining metric, I have four profiles defined which are showing the VM remaining. So the profiles take 4 VM profiles into consideration Small, Average, Medium & Large. Well this is better but not best because again, we are dependent on the VM profile sizes which are defined automatically by vROps.



Now with vROps 6.1, this has changed a little bit and for the best. Let me show you what I am talking about. If we look at vROps 6.0.x, we can see that the default 4 profiles here:

Select a Cluster from the Inventory List -> Click on Analysis Tab -> Click on Capacity Remaining and you would see the 4 profiles and the VMs remaining.


Here you can see that we have 4 profiles which are showing the virtual machines remaining in the cluster. If you click on the black triangle on any of the profiles, it would display the Profile Definition calculated automatically by vROps, see the screenshot below:

Remember, this is fixed and cannot be altered with any version of vROps except version 6.1. Yes, you heard it right, you can define custom profile sizes in vROps 6.1 which can match the standard offerings from your service catalog (if you have one) and make the results more predictable for the business in terms of how many more VMs and what type of more VMs before you actually run out of capacity. Read the above again. Yeah.. THIS IS SUPERCOOL because now I have better visibility into capacity remaining which resonates to my business and not just a random size.

Without further ado let's have a look at the same on vROps 6.1.


While you see the standard profiles here defined out of the box, you also have an opportunity to click on that box highlighted in RED and click on the + Sign to define a profile of your own. This is how you can get more predictive results which relate to your own environment. Let us click on the plus sign to add a new profile.



Let us fill this up and see what inputs go into creating this profile:



You can clearly see the options available to either model out of an existing VM or define your own profile. You even have the option to model anything which you are running capacity planning on, whether Virtual Machine, Datastores or 3rd party object. Finally you can chose between Allocation based Capacity Modelling or Demand based Capacity Modelling. Let's click on OK to save this and refresh the window to check if the new profile is available:


You can see that the new profile is available now, however the calculation on the number of VMs remaining shows a question mark. This is because the next Capacity Calculation would happen at 9 PM in the night. This is the default time of the vROps instance to run the capacity engine and redo all the capacity calculations every night.

So, wait for one more day and you would start seeing values there. Or wait for for my next post where I will post a tweak by which you can Force the Capacity engine to run on demand ;-) (GRIN..). Till then, enjoy the new feature and implement it to meet your business requirements.

And do not forget to SHARE & SPREAD THE KNOWLEDGE!!