Showing posts with label Architecting vSphere. Show all posts
Showing posts with label Architecting vSphere. Show all posts

Wednesday, February 5, 2014

Part 6 (Final) - Architecting vSphere for Business Critical Apps - A Scoop From my vForum Prezo!

This article throws light on Architecting vSphere Infrastructure for running workloads which are critical to your business. With this article I will end this series which I have been writing on Architecting vSphere Infrastructure. I thought I would conclude this with something which I am interested in evangelizing in the days to come as I not only see a huge learning & business opportunity in this area, but I also see the next major shift which virtualization is going to bring. 

Before we begin, here are the links to the first 5 articles in this series in case you are interested:







While creating my vForum Presentation (See Part 1 for more details), I realized that I would have a number of people in the crowd who will be interested in discussing the virtualization of their business critical workloads, since they have already exhausted all of their Tier 2 & Tier 3 applications running on physical x86 boxes, by converting them into virtual machines.

With this article we will look at some of the grey areas which we should be aware of when we are planning to Virtualize business critical applications on a vSphere Environment. At this point let me also highlight the fact that these are specific areas which may or may not suit all the applications/workloads. These are general recommendations which will help you iron out a lot of issues which I have seen organizations facing while taking this task on hand.

As always, let us first look at the single slide which I used during my presentation and then we will discuss each point as we move along.



While the above slide is quite self-explanatory, I would like to elaborate on a few of the pointers which might need some more explanation and clarification to put my point across to you and also due to recent changes in the vSphere 5.5 platform.

FOLLOW ASTBUL - This is not a technical jargon or a industry standard term. I coined this term to ensure that we do not rush into virtualization of business critical workloads. ASTBUL stands for ASSESS, SETUP, TEST, BENCHMARK, UAT & LIVE. When organizations begin with virtualization, they chose the low hanging fruits which are not only easy to virtualize, but can also live with downtime if it happens. We all know that it is not a rocket science to run a P2V conversion tool like converter to virtualize physical workloads. Similarly all the organizations with a decent virtualization footprint would have a Virtual First policy leading to virtualization of most of the new workloads as well. What does this mean to the C-Level people? Savings, infact let me say HUGE SAVINGS

While this was all good with Tier 2 and Tier 3 type of workloads, such an approach CANNOT be followed for applications/servers which are critical to an organization. What I mean to say is that you cannot, rather should not follow the approach of throwing business critical workloads on your virtual platform which is not architected or ready for them. You are actually making sure that all the hell will break lose and your whole idea of agility, savings, mobility etc with virtualization would be deemed foolish.

With BCA (Business Critical Applications) you need to ensure that you follow the strategy of ASTBUL. You begin with ASSESSMENT of the workload with tools like VMware Capacity Planner, SETUP a parallel to production setup in an isolated environment after inputs from application owners and hardware vendors, TEST the workloads in this environment to the maximum limit to see the Highs & Lows of compute, storage & network performance, BENCHMARK & publish the performance results to the business to get a buy in, run a User Acceptance Test (UAT) to get a buy in from the end users of the application and then finally pull the plug out of that EXPENSIVE physical server and  make your Virtual Workload LIVE. With this approach you will be successful 99% of the time. I believe in luck so would leave the rest of 1% on LUCK :-)


PLAY SMARTLY - The credit for the second point on the slide goes Michael Webster of Long White Virtual Cloud fame. Michael needs no introduction for the people who are active in the Virtualization Community, but for those who are not, he is the first VCDX in New Zealand and one of the finest craftsmen when it comes to Architecting vSphere Environments for running Business Critical Workloads. I would highly recommend you visit his blog to see all the great work he has done around Virtualization. Michael in one of his posts says that if you are virtualizing applications which are network intensive and they have a need to talk to other VMs in the same VLAN, one should try and place such VMs in an affinity. This will ensure that they will run on the same ESXi host and be a part of the same PortGroup/dvSwitch. In this way their network traffic would never leave the ESXi hosts and would transmit at the speed of memory between the VMs. This will not only save your network from choking but would also remove any performance bottlenecks which might result from sharing physical up-links. 

Their are a number of other tips & tricks which can help you get the best results for your business critical workloads. It is important that you TEST & BENCHMARK them as mentioned in the last point.


DE-MYSTIFYING NUMA - NUMA a.k.a. Non Unified Memory Access is another critical area to consider especially when you are virtualizing BCA. Their are a number of articles out their which discuss NUMA and I will not re-invent the wheel here by writing another one. Let's first look at the articles here and then I will add my 2 cents to this concept by trying to simplify & summarize this for you.

To begin with you should read this article from Wikipedia to understand what is NUMA. Then you should understand what is vNUMA and then read the following 3 articles:-




All the above mentioned articles are amazing as they clearly explain the concept of NUMA, vNUMA & the affects of assigning cores instead of sockets. The bottom line is this :-

  • Ensure that you do not cross the physical NUMA boundaries if possible. In-case you need to, GO AHEAD, but ensure that you have all the pre-requisites for vNUMA to do the smart scheduling for you.
  • Upto vSphere 5.1 vNUMA wont kick in if you use more than 1 core per socket while allocating vCPU AND you are crossing the NUMA boundary with your overall allocation. So do not use this setting unless you need to save on CPU license for the application and you can take the performance hit.
  • The point above nullifies with vSphere 5.5 as vNUMA in 5.5 is smart enough to ensure that it will allocate the number of vCPU required with the best possible combination of cores and sockets as per the architecture of the physical NUMA. Hence even if you goof up while allocation, vSphere will take care of it. For eg. if you assign 1 CPU with 16 cores, it will automatically size the machine as 2 CPU with 8 Cores each.


CRITICAL DOES NOT MEAN COMPLEX - When you speak to people about virtualizing business critical applications, the first thing which comes to their mind is "This has to be a complex architecture". This is where things start to become a little overwhelming as architects try to make use of each and every BEST PRACTICE and feature available. A classic example would be to run OS & Application clustering over and above vSphere HA. You need to go to the business and ask them the up-time requirement vs the cost of providing the up-time and you will always see that they are okay with the restart of HA rather than putting complex and expensive clustering technology which is difficult to deploy, manage and troubleshoot. Keep things as simple as possible & you will get the most optimum results from your efforts of running business critical applications in a virtual machine.


VIRTUALIZING BCA- IT'S NOT ABOUT THE MONEY HONEY - While the prime attraction towards virtualization was always monetary benefits, this mindset has to change with BCA (Business Critical Applications). In simpler terms, you might be saving money by moving your business critical workloads from a UNIX to LINUX platform and then virtualize it with vSphere, what you should NOT do is to practice things live OVER-COMMITMENT or HIGH CONSOLIDATION. All the money which you might be saving should be re-invested to ensure that you are successful at virtualizing business critical workloads. This investment should be done in running parallel Test Setups, Running Benchmarks, Zero Commitment of CPU & Memory, High Performance Storage platforms etc. Remember savings is not the only benefit of virtualization. Things such as Ease of Upgrade, scaling up of RAM and CPU on the fly, options to take Snapshots before Upgrading, quickly Rolling out new services, options to Clone from templates, easily moving from TEST to UAT to PROD and vice-versa, Zero-Downtime during hardware maintenance without having to invest in complex clustering solutions are some of the major benefits you get from virtualizing Business Critical Applications and this is what you should AIM for.


I would want to conclude by saying that it is an Art and not a Science, so its important that you INVOLVE THE EXPERTS. Its a fact that you need more experience than skills to architect a vSphere environment suitable for running business critical applications which is stable, secure & highly available.

With that note I will close this article and the entire series. I hope this has and will help you while you architect better for your organization or a customer. If you have any questions, comments or feedback, feel free to use the comments section and I would love to have a constructive debate on this and other topics.


Share & Spread the Knowledge!!



Friday, January 31, 2014

Part 5 - Architecting vSphere Networks: A Scoop From my vForum Prezo!

This is the fifth part of the series of articles I have been writing about “Architecting vSphere Environments”. For those who have not read the other parts, I would highly recommend you go through them to understand the context of this entire series and to pick up key learning's which I have gained during design & implementation of VMware vSphere Infrastructures for VMware customers. Here is the list:-






Maintaining the context from my previous articles, I would continue to stress on areas which are mostly ignored either due to high complexity or due to lack of skill sets. Also as the title of this article suggests, I will speak about key considerations around designing Networks for a vSphere Infrastructure. In my humble opinion, a single blog article cannot do justice to plethora of networking gotchas which you need to be aware of, hence I would expect that the reader of this article would have basic understanding of the concepts of Networking in a Virtual Environment. Needless to say that you should understand the traditional physical networks.

Let me make your’ s & mine task simpler by putting up a slide here which talks about 4 major things which you need to remember. These are my Do’s & Don’ts and are areas which are mostly mis-configured and mis-designed in most of the environments I have seen. Let’s have a look at the slide first and then we will discuss these options in detail.




















Network Layout

The first part on the slide speaks about laying out the vSphere Virtual Networks appropriately. The best and the worst part about doing this is that there is no right or a wrong way. This is one of the areas of vSphere which is entirely dependent on the Requirements & Constraints. While you need to ensure that you meet the requirements, you cannot step out of the boundaries created due to constraints. Let us see as to what are the key considerations you need to worry about.

  • One Role Per PortGroup – Using each port group for a single role would be an ideal scenario while designing networks for a vSphere environment. In essence, you should have one VMkernel (Management) PortGroup each for Management Network, Secondary Management Network (if available), FT, vMotion & vSphere Replication (this option is still not active in vSphere, although can be seen through vSphere Web Client). This will ensure that you assign individual VMK IDs to each function which not only makes your network settings simple but easy to troubleshoot as well. For Virtual Machine Port Groups you have to use separate PortGroups for easy identification and VLAN tagging. Although with a VLAN trunk option available on dvSwitch you can use a single PortGroup for multiple segments (rarely used).
  • Use of VLANs – Use of VLANs is not uncommon in a vSphere infrastructure. This is the easiest method to make sure that you can run multiple network segments on a single physical wire. I have seen a very few environments who do not use VLANs & in most of such cases it is the lack of networking knowledge. Have seen servers with up to 22 NIC ports and the same number of network cables going out of the server which ofcourse is INSANE. The simple reason behind this crazy setup was no use of VLANs. Last but not the least, don’t forget to trunk the VLANs in the physical switch ports and define the same on the virtual port group. Use of VLAN1 (native VLANs) is another crime which should never be committed :-)
  • Standardize VMkernel interfaces across ESXi hosts – This one is the simplest but the most ignored best practice of all. Each VMK defined on a host should be identical to the VMK defined on other hosts in the same ESXi cluster. For e.g. if VMK0 is Management on one host, it should be Management on all the other hosts in the cluster. Same applies to all the VMkernel interfaces. This will ensure that your network configuration is easily readable. Another benefit is seen when you are writing scripts to manage network settings of your ESXi cluster. Try to follow this practice for IP addressing of management functions as well. This makes your life easy when you are troubleshooting during those unplanned downtime.
  • Redundant NICs, dedicated FT Network, appropriate Load Balancing Policies & appropriate NIC failover ordering are other few configurations which can make or break the networking stack of a vSphere Infrastructure.


Enterprise Plus Licensing – Do you use a dvSwitch? 

If you are one of many who have Enterprise Plus licensing on vSphere, but are not using the Distributed Virtual Switch, you my friend are wasting your investments made on the vSphere Platform. The only reason for vSphere to be the number one hyper-visor in the world is the amazing features that comes with this platform. Distributed Virtual switch a.k.a. DVS is definitely one of them. In my experience every 6 out of 10 enterprise plus licensed customers I see are still running on VSS (Virtual Standard Switch). Although VSS might meet your requirements, it does not have the intelligence of a DVS. While managing an environment with a DVS is simple due to a centralized control plane, features like health check & backup/restore are priceless. Health Check ensures that any mis-configuration is automatically rolled back which reduces the risk of human errors which by the way is the single biggest reason for most of the issues you will see in any IT infrastructure. DVS also allows you to Backup itself and restore it in-case you need to replicate environments or re-build from scratch.

While I am pushing all my customers to dvSwitch, I would highly recommend you plan a move as well in-case you are still on the VSS. While you have a freedom of moving all the PortGroups to dvSwitch, some people prefer to keep the management networks (Heartbeat/FT/vMotion etc.) on VSS while they move the Virtual Machine PortGroups or Data traffic to DVS. I completely support this philosophy as well, especially when you do not have any redundancy for your vCenter Server. This allows you to play around with the network settings of an ESXi hosts by directly logging into it using the vSphere Client, even if the vCenter Server is down and not recoverable. Although it is a bad situation to be in, but certainly not unrecoverable.

Read the following knowledge base article from VMware which is a great source for clearly laying out the differences between Standard & Distributed Virtual Switches - Overview of vNetwork Distributed Switch concepts.


Convergence is the way forward - Move to 10G Networks

I know for a fact that many might challenge this thought. However, in my experience and a number of projects I have worked on, I see that customers with a converged infrastructure are more on ease and peace of mind as compared to the one who have their network all over the place.

While I say this, I would only encourage you to make this move if you have planned IT budgets, as this would be an overkill if you want to transform as in most of the cases your old hardware would be a waste. With 10G, you have options of either using the convergence with the management platform available from the hardware vendors or use features like Distributed Virtual Switches / Network I/O Control to manage the bandwidth requirements of your workload.

Here are a couple of articles which I wrote around using 10G cards for vSphere Networking. Would urge you to read them to see how things become simple with implementation of converged networks.



Some More Quick Tips Around vSphere Networking
  • Be a miser while provisioning vSwitches, Portgroups or for that matter any virtual hardware. Please remember they might appear to be FREE but they do use CPU Cycles and Memory since they are nothing but a piece of code which runs as soon as you create a new object.The idea is not to stop you from provisioning, but to make sure that you do not provision what you do not require. This also helps when you are trying to troubleshoot issues.
  • vMotion has been around for a while, but the options to configure vMotion have always been evolving. Use these options to ensure that you have Fast vMotion Networks which will allow you to load balance clusters, evacuate hosts & run maintenance tasks way faster then what you have been doing. Use Multi NIC vMotion to do this.  There are a number of articles on Multi NIC vMotion which you can find on frankdenneman.nl by Frank Denneman which I refer to. Go use them, they are awesome :-)
  • This is another important tip which is organizational and operational in nature. You need to involve your Networking team while choosing, designing & implementing your vSphere Infrastructure. The more the networking experts are involved, the better would be your networking stack supporting your virtual infrastructure. Also, as I see convergence in the Datacenter, I also see convergence in the IT teams. Networking & Storage teams are getting trained on vSphere and Wintel/VMware teams are getting trained on the networking and storage platforms. This change is welcoming and my suggestion to you would be follow this change ASAP so that you, your teams and your organization is able to cope with this paradigm shift.

With this, I will close this article. Hope this helps you make the right choices while choosing the network configuration for your infrastructure. Please share your comments, thoughts & feedback around this series.


Share & Spread the Knowledge!!


Follow on Twitter -     @Sunny_Dua

LinkedIn - LinkedIn@duasunny

Sunday, December 29, 2013

Part 4 - Architecting vSphere - Remember the Design Dimensions & Process - From My vForum Prezo!

This article is the Part 4 of the Series "Architecting vSphere Environments". Here are the other 3 parts which I would highly recommend to read:-

I have a very strong feeling that this should actually be the Part 1 of this entire series. In this article we will have a quick look at the different facets of a vSphere Design and also throw some light on the processes and procedures one needs to follow to create a successful architecture for a vSphere or that matter any IT infrastructure.

While I write this, I realize that processes can be a little boring as compared to technical stuff. However in my experience, no matter how technical one is, unless and until you have a correct approach to an architecture, you will end up failing 9 times out of 10. This approach to architecture forms the Process.

Let us have a quick look on some of the dimensions or facets which are involved in designing and architecting a vSphere Environment and briefly discuss them before looking at the stages in  design process. This time around I would like to give the content credit to VMware vSphere Design Book authors - Forbes Guthrie, Scott Lowe & Maish Saidel-Keesing. The way they explain this topic in the mentioned book is absolutely fantastic. The slide below depicts the same:-



I have seen articles about this concept from this book before. I still wanted to include this in my presentation and this article since I experience these facets on a daily basis while doing projects of  small to large scale. 

If you take up any environment which you need to architect, you would need to consider the Technical Facet, Organizational Facet & the Operational Facet. With each facet you need to ask yourself and the project members, questions would help you to create a design which not only help meet the requirements but also helps you define the scope of the project.

  • With Technical Dimension, you would go into your favorite questions which would usually be questions with "WHAT"?
  • With Organizational Dimension, you start looking at Responsibility, Authority and Accountability related concerns, hence the questions begin with "WHO"
  • &, With Operational Dimension, you look at the most important part of an Architecture which might impact the Operational Procedures and Processes.

Therefore, the above mentioned facets are very important and it is critical that we give them utmost importance in the entire process of Architecting a vSphere environment. With this, let's see what are the various stages in a design process:-



If you carefully walk through all the stages in the design process, you would end up with a successful Design and the right tools to implement and validate the design. With this I will close this article and I hope the recommendations in this article will help you adopt the right strategy when you architect a vSphere environment for any organization.


Share & Spread the Knowledge!!

Monday, December 16, 2013

Part 3 - Architecting Storage for vSphere Environments - A Scoop from my vForum Prezo!


In continuation to the series on "Architecting vSphere Environments", this post talks about Architecting the storage for a vSphere Infrastructure. For those who have not read the previous parts of this series, I would highly recommend you go through them in order to get a complete picture on the  considerations which matter the most while designing the various components of a vSphere Infrastructure.

Here are the links to the parts written before:-

Considering you have read the other parts, you would know that I have been talking about "Key Considerations" only in this series of articles, hence this post would be no different and would only talk about the most important points to keep in mind while designing the STORAGE architecture on which you will run your virtual machines. As mentioned before this comes from my experiences which I have gained from the field and from advises which I have read & discussed with a lot of Gurus in the VMware community.

With the changes in the storage arena in the past 2 to 3 years, I will not only talk about traditional storage design, but would also throw in some advice on the strategy of adopting Software Defined Storage (SDS). The slide below indicates the same.




I would start with talking about the Traditional Storage and the key areas. Let's begin with talking about IOPS (Input/Output Per Second). Have a look at the slide below.





The credit for the numbers and formula shown in the above slide goes to Duncan Epping. He has a great article which explains IOPS. Probably the first search result on Google if you search for the keyword "IOPS". I included this into my presentation because this is still the most miscalculated and ignored area in more than 50% of virtual infrastructures. In my experience only 2 out of 10 customers I meet discuss IOPS. Such facts worry me as deep inside I know that someday the Virtual Infrastructure would come down like a pack of cards if the storage is not sized appropriately. With this let's look at some of the key areas around IOPS.

  • Size for Performance & Not Capacity - Your storage array cannot be sized for the capacity of data which you need to store. You would always have to size for the Performance which you need. In 90% of the cases you would need to buy more disks than you need, in order to make sure that you meet the IOPS requirements of the workloads which you are planning to run on a Volume/LUN/Datastore.

  • Front End vs. Back End IOPS - This is the most common mistake which is committed while sizing storage. Though your intent might be correct to size the storage on the basis of workload requirements, please remember that workload demands FRONT-END IOPS while Storage Arrays provides BACK-END IOPS. In order to convert Front-end IOPS you need to consider Read/Write ratio of IOPS, the type of disk being used (i.e. SSD, SAS, SATA etc) and finally the RAID Penalty. I have explained this concept with the formula above (Courtesy - Duncan Epping)

An application owner asks you for 1000 IOPS for a workload which has 40% Reads and 60% Writes. These 1000 IOPS are front end IOPS. Here is how you will determine the Back-end IOPS on the basis of which you will architect the Volume/Datastore or may be buy disks in case you are into the procurement cycles.

Back-End IOPS = (Front-End IOPS X % READ) + ((Front-End IOPS X % Write) X RAID Penalty)

The RAID Penalty is shown in the slide above along with the number of IOPS which you receive with different types of disks available for a storage array today. So considering our example, if we are choosing RAID 5 for this workload, the raid penalty would be 4 IOPS. Let's do the math now:-


Back-End IOPS = (1000 X 40%) + ((1000 X 60%) X 4)
                           = (1000 X 0.4) + ((1000 X 0.6) X 4)
                           = (400) + (600 X 4)
                           = 400 + 2400
                           = 2800 IOPS

So you can Clearly see that 1000 Front-End IOPS mean 2800 Back-End IOPS. That is 2.8 times the actual requirement. While you can get 1000 IOPS from  just 7, 15k RPM SAS drives, you need a whopping 17 Disks to get 100 Front-End IOPS.

I hope that gives you a clear picture on the fact that you not only need to size for performance, you also need to size correctly for performance, since storage can make or break your virtual infrastructure. If you don't believe ask the VMware Technical Support team the next time you speak to them. A whopping 80% of the performance issue case which they deal with are related to a poorly designed storage.


Along with IOPS let us see a few more areas where we need to be cautious.



  • DO NOT simply upgrade from an older version of VMFS to a new version. Please note that I am referring to major releases only. In-fact to be more precise I am referring to an upgrade of VMFS 3.X to VMFS 5.X. For those who follow VMFS (Proprietary Virtual Machine File System developed by VMware) would know that VMFS 5.x was introduced with vSphere 5.0. Prior to this the version of VMFS was 3.x. VMware has done some major changes to VMFS 5.0 which changes the way the blocks and metadata functions. If you are upgrading from vSphere 4.1 or before to vSphere 5.x, then I would highly recommend to re-format the datastores and create them afresh to get the latest version of VMFS. An in-place upgrade from VMFS 3.x to 5.x would bring in the legacy features of the file-system and that can have a performance impact on operations such as vMotion & SCSI locking. You should empty each datastore by using Storage vMotion format it and then move the workloads back on it. Time consuming but absolutely worth it.
  • If you are using IP based storage, then you should have a separate IP fabric for the transport of storage data. This ensures security due to isolation and high performance since you have a dedicated network card / switching gear for storage. This is a very basic recommendation and people tend to overlook it. Please invest here and you would see that IP storage working at par with FC Storage.
  • Storage DRS is cool but should be partially used if you have a storage with Auto-Tiered disks. Auto-Tiering is usually available in the newer arrays and allows workloads to access HOT data (frequently used) off the fastest disks (SSD), while the cool data (not frequently used) is staged off to slower spindles such as SAS and SATA. The IO metric feature available within Storage DRS is made to improve the performance for NON-AUTO Tiered storage only. Hence, if you have an Auto-Tiered Array, please disable this feature since it will not do any good and simply become an overhead. Having said that, SDRS also gives you the feature of automatically placing virtual machines when they are first created on an appropriate datastore on the basis of capacity and performance requirements (with storage profiles). Hence use the Auto Tier feature of your storage to manage performance while use SDRS only for initial placement of VM.
  • With vSphere 5.5 the size of VMDKs and RDMs have been increased. They are now at 62 TB limit. Use them if you have monster workloads in your infrastructure and only if you need them. Over-provisioning in a VMware environment is as bad as under provisioning.
  • IOPS help you size, but throughput and multipathing calculations are also critical since they are the bridge between the source (servers) & the destination (Spindles). make sure you have the correct path policies and appropriate amount of throughput, qdepth for smooth transmission of data packets across fabric.
  • Coming to the last point, it is important to understand the nuances of the VMFS file system. Thin vs Thick disks, Thin on Thick or Thick on Thick etc are various considerations which you need to keep in mind. I would recommend you read this excellent VMware blog article written by Cormac Hogan, which talks about almost all the possible file system combinations which one can have. My 2 cents on this would be to keep things simple and to use EAGERED ZEROED THICK VMDKs for all your latency sensitive virtual machines. This will save you that extra time which ESXi takes to zero down the blocks before it could write new data in case of LAZY ZEROED THICK VMDKs. You do not have to panic in case you are reading this now and already have your latency sensitive workloads running on Lazy Zeroed Disks. There are a number of ways by which you can convert the existing Lazy Zeroed Disks to Eagered Zero. The slide below shows all the options.



Finally let's quickly move on to the things from the New World of Software Defined everything. Yesssss.. I am talking about Storage of the new era a.k.a. SOFTWARE DEFINED STORAGE (SDS). If you have a good memory, you would remember that I told you to size your storage for Performance and not Capacity. I am taking back my words right away and would want to tell you that with Software Defined Storage, you no longer have to Size for Performance. You just need to Size for Capacity and Availability. 

Let's have a look at the slide below and then we will discuss the key areas of software defined storage as conceptualized by VMware.



The above slide clearly indicates the 2 basic solutions from VMware in the arena of Software Defined Storage. The first is vFRC (vSphere Flash Read Cache) which allows you to pool cache either from SSD or PCI Cards as the first read device, hence improving read performance for read intensive workloads such as Collaboration Apps, Databases etc. Imagine this as ripping of the cache from your storage array and bringing it close to the ESXi server. This cuts down the travel path and hence improves the performance incredibly.

Another feather in the cap of Software Defined Storage is VSAN. As Rawlinson of Punching Clouds fame always says, it's VSAN not with a small "v" like other VMware products. Only he knows the secret behind the uppercase "V" of VSAN. Earlier this year I met Rawlinson at a technology event in Malaysia and he gave me some great insights on VSAN. 

VSAN uses the local disks installed on an ESXi server which has to be combination of SSD + SAS or SATA drives. The entire storage from all the ESXi hosts in a cluster (currently the limit is minimum 3 and maximum 8 nodes) is pooled together in 2 tiers. The SSD tier for Performance and the SAS/SATA tier for capacity. To read more about VSAN I would suggest you to read a series of articles written by Duncan available on this link and Cormac on this link.

Remember VSAN is in beta right now, hence you can look at trying it for the use cases which I have mentioned in the slide above and not for your core production machines. 

With this I will close this article and I hope this will guide you to make the right choices for choosing, architecting and using storage be it traditional or modern for your vSphere Environments. feel free to comment and share your thoughts and opinions around this post.


Share & Spread the Knowledge!!



Tuesday, November 26, 2013

Part 1 - Architecting vSphere Clusters - A scoop from my vForum Prezo!

A few days back, I got an opportunity to present at vForum 2013 in Mumbai, the Financial Capital of India. With more than 3000 participants across 2 days of this mega event, it was definitely one of the biggest customer events in India. I along with my team was re-presenting the VMware Professional Services at vForum and I was given the opportunity present on the following topic:-

"Architecting vSphere Environments - Everything you wanted to know!"

When we finalized the topic, I realized that the presenting this topic in 45 minutes is next to impossible. With the amount of complexity which goes into Architecting a vSphere Environment, one could actually write an entire book. However, the task on hand was to narrate the same in form of a presentation. 

As I started planning the slides, I decided to look at the architectural decisions, which in my experience are the "Most Important One's". These are important decisions as they can make or break the Virtual Infrastructure. The other filtering criterion was to ensure that I talk about the GREY AREAS where I always see uncertainty. This uncertainty can transform a Good Design into a Bad design. At the end I was able to come out with a final presentation which was received very well by the attendees. I thought of sharing the content with the entire community through this blog series and this being the Part 1, where I will give you some key design considerations for designing vSphere Clusters.

Before I begin, I would also want to give the credit to a number of VMware experts in the community. Their books, blogs and the discussions which I have had with them in the past, helped me in creating this content. This includes books & blogs by DuncanFrankForbes GuthrieScott LoweCormac Hogan & some fantastic discussions with Michael Webster earlier this year.


Here is a small Graphical Disclaimer:-



Here are my thoughts on creating vSphere Clusters!!


The message behind the slide above is to create vSphere Clusters based on the purpose they need to fulfill in the IT landscape of your organization.

Management Cluster

The management cluster refers here to a 2 to 3 host ESXi host which is used by the IT team to primarily host all the workloads which are used to build up a vSphere Infrastructure. This includes VMs such as vCenter Server, Database Server, vCOps, SRM, vSphere Replication Appliance, VMA Appliance, Chargeback Manager etc. This cluster can also host other infrastructure components such as Active Directory, Backup Servers, Anti-virus etc. This approach has multiple benefits such as:-

  • Security due to isolation of management workloads from production workloads. This gives a complete control to the IT team on the workloads which are critical to manage the environment.
  • Ease of upgrading the vSphere Environment and related components without impacting the production workloads.
  • Ease of troubleshooting issues within these components since the resources such as compute, storage and network are isolated and dedicated for this cluster.


A quick tip would be to ensure that this cluster is minimum a 2 node cluster for vSphere HA to protect workloads in case one host goes down. A three(3) node management cluster would be ideal since you would have the option of running maintenance tasks on ESXi servers without having to disable HA. You might want to consider using VSAN for this infrastructure as this is the primary use case which both Rawlinson & Cormac suggest. Remember, VSAN is in beta right now, so make your choices accordingly.


Production Clusters

As the name suggests this cluster would host all your production workloads. This cluster is the heart of your organization as this hosts the business applications, databases, web services, literally this is what gives you the job of being a VMware architect or a Virtualization Admin. J

Here are a few pointers which you need to keep in mind while creating Production Clusters:-


  • The number of ESXi hosts in a cluster will impact you consolidation ratios in most of the cases. As a rule of thumb, you will always consider one ESXi host in a 4 node cluster for HA failover (assuming), but you could also do the same on a 8 node cluster, which ideally saves 1 ESXi host for you for running additional workloads. Yes, the HA calculations matter and they can be either on the basis of slot size or percentage of resources.
  • Always consider at least 1 host as a failover limit per 8 to 10 ESXi servers. So in a 16 node cluster, do not stick with only 1 host for failover, look for at least taking this number to 2. This is to ensure that you cover the risk as much as possible by providing additional node for failover scenarios
  • Setting up large clusters comes with their benefits such as higher consolidation ratios etc., they might have a downside as well if you do not have the enterprise class or rightly sized storage in your infrastructure. Remember, if a Datastore is presented to a 16 Node or a 32 Node cluster, and on top of that, if the VMs on that datastore are spread across the cluster, chances that you might get into contention for SCSI locking. If you are using VAAI this will be reduced by ATS, however try to start with small and grow gradually to see if your storage behavior is not being impacted.



·    Having separate ESXI servers for DMZ workloads is OLD SCHOOL. This was done to create physical boundaries between servers. This practice is a true burden which is carried over from physical world to virtual. It’s time to shed that load and make use of mature technologies such as VLANs to create logical isolation zones between internal and external networks. In worst case, you might want to use separate network cards and physical network fabric but you can still run on the same ESXi server which gives you better consolidation ratios and ensures the level of security which is required in an enterprise.


Island Clusters

Yes they sound fancy but the concept of Island clusters as laid down in my slides is to run islands of ESXi servers (small groups) which can host workloads which have special license requirements. Although I do not appreciate how some vendors try to apply illogical licensing policies on their applications, middle-ware and databases, this is a great way of avoiding all the hustle and bustle which is created by sales folks. Some of the examples for Island Clusters would include

·    Running Oracle Databases/Middleware/Applications on their dedicated clusters. This will not only ensure that you are able to consolidate more and more on a small cluster of ESXi hosts and save money but also ensure that you ZIP the mouth of your friendly sales guy by being in what they think is License Compliance.

·    I have customers who have used island clusters of operating systems such as Windows. This also helps you save on those datacenter, enterprise or standard editions of Windows OS.

·    Another important benefit of this approach is that it helps ESXi use the memory management technique of Transparent Page Sharing (TPS) more efficiently since with this approach there are chances that you are running a lot of duplicate pages spawned by these VMs in the physical memory of your ESXi servers. I have seen this going up-to 30 percent and this can be fetched in a vCenter Operations Manager report if you have that installed in your Virtual Infrastructure.

With this I would close this article. I was hoping to give you a quick scoop in all these parts, but this article is now four pages J. I hope this helps you make the right choices for your virtual infrastructure when it comes to vSphere Clusters.

Stay tuned for the other parts in the near future…


As always – Share & Spread the Knowledge!!