Amazon Web Services – Elastic VMware Service – A Few Notes …

Amazon Web Services (AWS) recently announced the general availability of its Elastic VMware Service (EVS). This is a brief post covering what the service is, how it can be consumed, and why it might be of interest.

 

What Is It?

EVS is a service that provides the ability to run VMware Cloud Foundation (5.2.1) on AWS bare-metal instances (specifically I4i.metal). It’s a by-the-book consolidated architecture (running management and workload domains together), and contains the following VCF management components:

  • ESXi hosts;
  • vCenter Server instance;
  • SDDC Manager;
  • vSAN datastore;
  • Three-node NSX Manager cluster;
  • vSphere cluster; and
  • NSX Edge cluster.

Storage is provided by vSAN (and selected third-party storage integrations), with two discrete network layers being delivered: the Amazon VPC and the VMware NSX overlay network. You can read more about the architecture here.

To consume the service, you’ll need 4 nodes to get started, along with at least 2 VPC Route Servers, and AWS Enterprise Support. There’s a pricing calculator on the AWS website. You’ll also need to procure VCF licensing from Broadcom or an authorised reseller.

 

Where Is It Available?

The initial launch Regions are:

  • US East (North Virginia)
  • US East (Ohio)
  • US West (Oregon)
  • Asia Pacific (Tokyo)
  • Europe (Frankfurt)
  • Europe (Ireland)

Availability in more Regions is planned for the future.

Other Considerations

Storage

You can leverage NetApp’s FSx for NetApp ONTAP service to provide additional storage for the cluster. There’s some useful guidance here and here. If you like it more orange, Pure Storage’s Cloud Block Store is also available for use with EVS. You can read that announcement here.

Clusters

Cluster sizes are currently limited to between 4 (minimum) and 16 (maximum) nodes. There is currently no support for stretched clusters. This will likely change in the future.

Add-ons

Services that VMware Cloud on AWS users may be familiar with, such as vDefend Advanced Threat Protection, and VMware Live Cyber Recovery, are not currently available with EVS.

Planning and Deployment

Much like VMC on AWS (and most other cloud service deployments, for that matter), there’s a bit of planning you’ll need to do before you jump on the console. There’s a handy checklist that takes you through all of the requirements here. It’s also worth noting that, unlike VMC on aWS, this isn’t a managed service, so some of the VPC constructs you see with VMC on AWS aren’t used with EVS. For example, there’s no shadow VPC deployed for EVS – it’s all done out of the customer account. And features like EDRS aren’t available, you need to manage your capacity yourself, in much the same way as you would do with an on-premises VCF deployment.

 

Thoughts

EVS is the same as VMC on AWS in the sense that it’s running VMware software on top of AWS bare metal servers. There are plenty of differences though, some of which I’ve covered above. I imagine many of the constraints around additional product integrations and scalability will be removed in the near future as the product evolves, and we’ll no doubt see some changes when it comes to partners offering managed services wrapped in with EVS. It’s been a while coming, but now that it’s here, I’m interested to see how the market responds. There are probably plenty of questions I haven’t answered, but there’s a useful FAQ document here.

Brisbane VMUG – July 2025

The July 2025 edition of the Brisbane VMUG meeting will be held on Thursday 24th July at the Pig ‘N’ Whistle Riverside, Riverside Centre, 123 Eagle St, Brisbane City QLD 4000 from 3pm to 5pm. It’s sponsored by Druva and promises to be a great afternoon. Register here.

Here’s the agenda:

Session 1 (45min – 1hr): Cyber Recovery for VMware (Druva)

Presenter: Bruce Perram (Senior Systems Engineer)

Topic: Many organisations today are focused on prevention of cyber-attacks however organisations are still being breached. How do you recover from a cyber-attack? This session will include a demo of cyber recovery in a VMware environment. It will also explore the difference between backup & recovery from cyber recovery and the capability required rapidly perform a cyber recovery.

Session 2 (45min – 1hr): VCF 9 – The High Points and Demo (Broadcom)

Presenters: Peter Hauck (Staff Solutions Architect) and David Miller (Senior Solutions Architect)

Topic: With VCF 9 release last month there is a lot of new and exciting features to show. We will leave the details feature discussions to the excellent marketing coming from Broadcom HQ. In this session we run through a live demo of the solution and discuss what it means for customers today. The session is live so lots of opportunity to ask questions.

Druva has gone to great lengths to make sure this will be a fun and informative session. Following the sessions, the VMUG team will be hosting vBeers – an opportunity to network with other members of the community and chat with the experts!

Exciting giveaways await! You could score prizes like a VMUG Advantage subscription or a large LEGO kit – just make sure to register to be eligible.

Random Short Take #103

Random short take #103. Want news? You’re probably not in the right place. Let’s get random.

VMware Cloud on AWS – TMCHAM – Part 14 – Stretched Clusters and Disaster Recovery

In this edition of Things My Customers Have Asked Me (TMCHAM), I’m going to cover some of the things to consider when looking at Disaster Recovery (DR) and availability solutions on the VMware-managed VMware Cloud on AWS platform. It might seem like basic stuff if you’re already using the platform, but there are plenty of customers I talk to that are struggling to distinguish between DR, data availability, and periodic data protection in a VMware Cloud context.

 

Terminology

I’m cheating a little here and re-using some of my older content, but I think it still holds up.

  • Disaster Recovery – Disaster Recovery is the ability to recover applications after a major event (think flood, fire, data centre is now a hole in the ground). This normally involves a failover of workloads from one DC to another in an orchestrated fashion. In a VMware context, you would use VMware Live Recovery (Site or Cyber, depending on the environment).
  • Disaster Avoidance – Disaster avoidance “is an anticipatory strategy that is in place in order to prevent any such instance of data breach or losses. It is a defensive, proactive approach to keeping data safe” (I’m quoting this from a great blog post on the topic here). In a VMware context, this would be a stretched cluster using vSAN on VMware Cloud Foundation or in VMC on AWS.
  • Periodic Data Protection – This is the kind of data protection activity we normally associate with “backups”. It is usually a daily activity (or perhaps as frequent as hourly) and the data is normally used for ad-hoc data file recovery requests. Some people use their backup data as an archive. They’re bad people and shouldn’t be trusted. You should consider PDP as separate to DA or DR solutions.

 

So How Do I Best Protect My Workloads? (Part 1)

Understand Your Workloads

You need to understand how your applications work. And you need to understand what talks to them and what they talk to. There are plenty of ways of doing this, including a variety if applications available that scan network traffic and draw pretty pictures of data flows for you. From experience, there’s a very good chance you won’t get this information by reading current documentation, because it won’t be current.

You need to understand how important your applications are to you. Do you need to recover them immediately if something goes wrong? Can you do without them for a few hours? A few days? How much data can you afford to lose? All of it? (in which case why do you still have this data sitting around?). None of it? Some of it? You need to be disciplined and thorough in your assessment of your workloads and what they mean to your business, because protecting them or recovering them can be an expensive exercise.

Understand What Scares You Most

What do you call a disaster? Is it a data centre turning into a hole in the ground? Do you run most of your workloads in an area prone to natural disasters? Or are you more worried about people in hoodies breaking into your infrastructure and encrypting it and your data protection infrastructure? (Hot tip: you absolutely should be worried about this). You need to be fairly paranoid about the scenarios you’re trying to protect your environment from. No one thought it was a problem keeping workloads in two buildings in New York until it was. At the same time, you’ll need to introduce some pragmatism into your planning. I’ve had customers ask me what they could do if all of the Sydney Region was wiped out (not just the AWS deployment, but the whole city). Could they recover data to Melbourne? Sure. Would they have any staff left to do it? Or want to do it? Depending on the nature of the disaster, the technical aspects of disaster recovery are going to be the easy bits. You still need to make sure the people and process bit is covered.

Who Can Help You? 

Do you want to handle any recovery activities yourself when things go bang? Or are you happy to let someone else do the heavy lifting for you? Some folks don’t trust outsiders to do “important” things, like recover workloads, whilst other organisations would consume absolutely everything as a service if it was an option.

How Much Can You Spend?

You need to answer the questions above before you can answer how much you’re willing to spend on one (or more) solutions to help keep your workloads protected. If your apps aren’t that important, or your business can continue to run without access to these apps for a period of time, then you’re going to be less inclined to go for the “platinum” solution. But if you’re whole business relies on applications being available 24/7/365, then you’re hopefully more understanding that you’ll need to invest.

 

So How Do I Best Protect My Workloads? (Part 2)

Data Availability with VMware Cloud on AWS

As I mentioned earlier in this post, if you’re looking to provide data availability for your workloads on VMware Cloud on AWS, you should be looking at deploying a stretched cluster solution. In short, stretched clusters in VMware Cloud on AWS are designed to protect against an AWS availability zone (AZ) failure. There’s a feature brief available here that goes into more depth on the solution. Key things to note are that the vSAN stretched cluster is deployed across 3 AZs (2 workload, 1 witness). So you could have 3 hosts in one AZ, 3 in another, and VMware deploys a witness node in the third site as part of the deployment. From the document, “vSAN Fault Domains are configured to inform vSphere and vCenter which Hosts reside in which AZ. Each Fault Domain is named after the AZ it resides within to increase clarity. Logical networks are also extended using NSX to support workload mobility across the two AWS availability zones”. You also have the option to keep data using various vSAN storage policies, including:

  • Dual-site mirroring (stretched cluster)
  • None – Keep data on primary (stretched cluster)
  • None – Keep data on secondary (stretched cluster)

Why would you not just mirror everything? Cost. Some applications simply don’t need that level of resilience, but it’s too complicated to deploy multiple SDDCs to support multiple deployment topologies (you can’t have a mix of single AZ and stretched clusters in the same SDDC), so customers can keep a mix of workloads in the same deployment. Note also that external storage solutions (like FSx for NetApp ONTAP) aren’t supported for use with stretched clusters. vSphere HA will restart the VM on the surviving AZ if a site failure is detected. There’s a short video on YouTube that covers a stretched cluster deployment and is worth looking at.

Disaster Recovery with VMware Cloud on AWS

I’ve covered VMware Live Cyber Recovery (and its predecessor, VMware Cloud Disaster Recovery) on this blog before. It’s a great solution to protect on-premises, GCVE, and VMware Cloud on AWS workloads.  It provides protection to a secure cloud filesystem with a Recovery Point Objective as low as 15 minutes, the ability to perform guided ransomware recovery in an isolated environment, and SaaS-based recovery automation. The kind of things you want to consider when evaluating the suitability of VLCR to your needs would be:

  • Recovery Point Objective – the maximum amount of time in which data may have been permanently lost during an incident. VLCR goes down to 15 minutes, but if you need lower than that you need to look at other architectures.
  • Scalability – how many Virtual Machines and how much storage are you trying to protect? And how quickly do you want to recover it? And where do you want to recover it to? Some folks don’t like VMware Cloud on AWS being the only recovery option, but VMware has announced “plans to enhance VMware Live Recovery support for Google Cloud VMware Engine (GCVE) as a cyber and disaster recovery (DR) site, for both on-premises as well as GCVE environments“. You can check out the configuration maximums for VMware solutions here.
  • Networking and security – how will you access your recovery workloads? Will you use HCX to have stretched networks between your protected sites and your recovery site? What does that look like for your core networking infrastructure? Is your CISO happy to recover stuff onto a cloud platform? There are many things you’ll need to work through internally to ensure people are comfortable. I’d also recommend you go with a pilot light deployment as opposed to an on-demand solution. While the pilot light option seems more expensive, it will save you when something goes wrong and you’ve already tested your deployment.
  • Other Considerations – there’s a bunch of info in the legacy VCDR Release Notes that highlight some of the constraints to be aware of from an architectural standpoint, and understanding some of these (like the ability to protect stretched clusters, but not recover to stretched clusters) is key to ensuring your architecture will work when you push the button. And one point that is often missed by folks looking at VLCR for the first time is that, as the recovery SDDC is just a VMware Cloud on AWS deployment, you can use it for other workloads in the meantime, like dev/test stuff, when you’re not in a recovery situation.

Periodic Data Protection with VMware Cloud on AWS

There are a tonne of different things you can do here, and I’d be naive to suggest that I could cover them all effectively in this post. To use periodic data protection solutions with VMware Cloud on AWS, you should check out this KB article, and work with your data protection vendor to ensure that the solution has been tested and works on VMware Cloud on AWS. I’ve encountered a few situations where the vendor has assumed that VMware Cloud on AWS works the same way as VMware on-premises, and with the same ability to do silly things with privileged access to ESXi hosts and things of that nature. There aren’t a lot of restrictions in place, but there are some important ones, like the aforementioned limits on direct host access, restricted account privileges, and so on. Just make sure you test it before you buy it. I’ll try and cover some of the scenarios that do and don’t work in a future post (without naming too many names).

 

Conclusion

I’m sticking with solutions available with VMware Cloud on AWS in this post, mainly because that’s what I’m talking to people about on a daily basis. But there’s plenty here that can be applied more generically to a variety of infrastructure solutions. You’ll have the same questions you’ll need to ask when it comes to protecting workloads on-premises (regardless of your choice of hypervisor), and you’ll need to make some hard decisions ab out what you want to protect depending on the value of what your workloads deliver to the business, and the available budget. You might also need to consider doing a combination of all three options, depending on the importance of the data / workloads you’re looking after. I’d always pick stretched clusters to keep my workloads up and available, but that won’t help me if some hacker decides to dump ransomware in my environment and encrypt all of my stuff. But if I need to go back 24 months to retrieve some file for an auditor, I wouldn’t use VMware Live Cyber Recovery for that. I should also mention that you should be using application-level protection as your first option for application resilience too. This other stuff is all infrastructure level, and potentially not as elegant, depending on your applications. But that doesn’t mean you can’t combine both methods. Unfortunately, there’s no one magical fix for this problem, but with a bit of planning, and some critical evaluation, you can protect your workloads against all manner of bad situations. And if you can’t do everything, at least do something, but just make sure people in your business understand the difference between the two. And remember, you’re only as good as your last recovery.

Random Short Take #60

Welcome to Random Short take #60.

  • VMware Cloud Director 10.3 went GA recently, and this post will point you in the right direction when it comes to planning the upgrade process.
  • Speaking of VMware products hitting GA, VMware Cloud Foundation 4.3 became available about a week ago. You can read more about that here.
  • My friend Tony knows a bit about NSX-T, and certificates, so when he bumped into an issue with NSX-T and certificates in his lab, it was no big deal to come up with the fix.
  • Here’s everything you wanted to know about creating an external bootable disk for use with macOS 11 and 12 but were too afraid to ask.
  • I haven’t talked to the good folks at StarWind in a while (I miss you Max!), but this article on the new All-NVMe StarWind Backup Appliance by Paolo made for some interesting reading.
  • I loved this article from Chin-Fah on storage fear, uncertainty, and doubt (FUD). I’ve seen a fair bit of it slung about having been a customer and partner of some big storage vendors over the years.
  • This whitepaper from Preston on some of the challenges with data protection and long-term retention is brilliant and well worth the read.
  • Finally, I don’t know how I came across this article on hacking Playstation 2 machines, but here you go. Worth a read if only for the labels on some of the discs.

Random Short Take #56

Welcome to Random Short Take #56. Only three players have worn 56 in the NBA. I may need to come up with a new bit of trivia. Let’s get random.

  • Are we nearing the end of blade servers? I’d hoped the answer was yes, but it’s not that simple, sadly. It’s not that I hate them, exactly. I bought blade servers from Dell when they first sold them. But they can present challenges.
  • 22dot6 emerged from stealth mode recently. I had the opportunity to talk to them and I’ll post something soon about that. In the meantime, this post from Mellor covers it pretty well.
  • It may be a Northern Hemisphere reference that I don’t quite understand, but Retrospect is running a “Dads and Grads” promotion offering 90 days of free backup subscriptions. Worth checking out if you don’t have something in place to protect your desktop.
  • Running VMware Cloud Foundation and want to stretch your vSAN cluster across two sites? Tony has you covered.
  • The site name in VMware Cloud Director can look a bit ugly. Steve O gives you the skinny on how to change it.
  • Pure//Accelerate happened recently / is still happening, and there was a bit of news from the event, including the new and improved Pure1 Digital Experience. As a former Pure1 user I can say this was a big part of the reason why I liked using Pure Storage.
  • Speaking of press releases, this one from PDI and its investment intentions caught my eye. It’s always good to see companies willing to spend a bit of cash to make progress.
  • I stumbled across Oxide on Twitter and fell for the aesthetic and design principles. Then I read some of the articles on the blog and got even more interested. Worth checking out. And I’ll be keen to see just how it goes for the company.

*Bonus Round*

I was recently on the Restore it All podcast with W. Curtis Preston and Prasanna Malaiyandi. It was a lot of fun as always, despite the fact that we talked about something that’s a pretty scary subject (data (centre) loss). No, I’m not a DC manager in real life, but I do have responsibility for what goes into our DC so I sort of am. Don’t forget there’s a discount code for the book in the podcast too.

Random Short Take #55

Welcome to Random Short Take #55. A few players have worn 55 in the NBA. I wore some Mutombo sneakers in high school, and I enjoy watching Duncan Robinson light it up for the Heat. My favourite ever to wear 55 was “White Chocolate” Jason Williams. Let’s get random.

  • This article from my friend Max around Intel Optane and VMware Cloud Foundation provided some excellent insights.
  • Speaking of friends writing about VMware Cloud Foundation, this first part of a 4-part series from Vaughn makes a compelling case for VCF on FlashStack. Sure, he gets paid to say nice things about the company he works for, but there is plenty of info in here that makes a lot of sense if you’re evaluating which hardware platform pairs well with VCF.
  • Speaking of VMware, if you’re a VCD shop using NSX-V, it’s time to move on to NSX-T. This article from VMware has the skinny.
  • You want an open source version of BMC? Fine, you got it. Who would have thought securing BMC would be a thing? (Yes, I know it should be)
  • Stuff happens, hard drives fail. Backblaze recently published its drive stats report for Q1. You can read about that here.
  • Speaking of drives, check out this article from Netflix on its Netflix Drive product. I find it amusing that I get more value from Netflix’s tech blog than I do its streaming service, particularly when one is free.
  • The people in my office laugh nervously when I say I hate being in meetings where people feel the need to whiteboard. It’s not that I think whiteboard sessions can’t be valuable, but oftentimes the information on those whiteboards should be documented somewhere and easy to bring up on a screen. But if you find yourself in a lot of meetings and need to start drawing pictures about new concepts or whatever, this article might be of some use.
  • Speaking of office things not directly related to tech, this article from Preston de Guise on interruptions was typically insightful. I loved the “Got a minute?” reference too.

 

Random Short Take #32

Welcome to Random Short Take #32. Lot of good players have worn 32 in the NBA. I’m a big fan of Magic Johnson, but honourable mentions go to Jimmer Fredette and Blake Griffin. It’s a bit of a weird time around the world at the moment, but let’s get to it.

  • Veeam 10 was finally announced a little while ago and is now available for deployment. I work for a service provider, and we use Veeam, so this article from Anthony was just what I was after. There’s a What’s New article from Veeam you can view here too.
  • I like charts, and I like Apple laptops, so this chart was a real treat. The lack of ports is nice to look at, I guess, but carrying a bag of dongles around with me is a bit of a pain.
  • VMware recently made some big announcements around vSphere 7, amongst other things. Ather Beg did a great job of breaking down the important bits. If you like to watch videos, this series from VMware’s recent presentations at Tech Field Day 21 is extremely informative.
  • Speaking of VMware Cloud Foundation, Cormac Hogan recently wrote a great article on getting started with VCF 4.0. If you’re new to VCF – this is a great resource.
  • Leaseweb Global recently announced the availability of 2nd Generation AMD EPYC powered hosts as part of its offering. I had a chance to speak with Mathijs Heikamph about it a little while ago. One of the most interesting things he said, when I questioned him about the market appetite for dedicated servers, was “[t]here’s no beating a dedicated server when you know the workload”. You can read the press release here.
  • This article is just … ugh. I used to feel a little sorry for businesses being disrupted by new technologies. My sympathy is rapidly diminishing though.
  • There’s a whole bunch of misinformation on the Internet about COVID-19 at the moment, but sometimes a useful nugget pops up. This article from Kieren McCarthy over at El Reg delivers some great tips on working from home – something more and more of us (at least in the tech industry) are doing right now. It’s not all about having a great webcam or killer standup desk.
  • Speaking of things to do when you’re working at home, JB posted a handy note on what he’s doing when it comes to lifting weights and getting in some regular exercise. I’ve been using this opportunity to get back into garage weights, but apparently it’s important to lift stuff more than once a month.