Pure//Accelerate 2018 – Wednesday General Session – Rough Notes

Disclaimer: I recently attended Pure//Accelerate 2018.  My flights, accommodation and conference pass were paid for by Pure Storage via the Analysts and Influencers program. There is no requirement for me to blog about any of the content presented and I am not compensated in any way for my time at the event.  Some materials presented were discussed under NDA and don’t form part of my blog posts, but could influence future discussions.

Here are my rough notes from Wednesday’s General Session at Pure//Accelerate 2018.

[A song plays. “NVMe” to the tune of Naughty By Nature’s OPP]

 

Charlie Giancarlo

Charlie Giancarlo (Pure Storage CEO) takes the stage. We share a mission: to power innovation. Storage is really important part of making that mission happen. We’re in the zettabyte era, we don’t even talk about EB anymore. It’s not the silicon age, or the Internet age, or the social age. We’re talking about the original gold rush of 1849. The amount gold in data is unlimited. We need the tools to pick the gold out of the data. The data heroes are us. And we’re announcing a bunch of tools to mine the gold.

Who am I? What has allowed us to get to success? Where are we going?

I’m new here, I’ve just gotten through his 3rd quarter. I’ve been an nngineer, entrepreneur, CTO, equity partner – entirely tech focused. I’ve made a living looking at innovation on the basis of looking at it as a 3-legged stool.

What are the components that advance tech?

  • Networking
  • Compute (Processing)
  • Storage

They advance on their own timelines, and don’t always keep pace, and the industry someitmes gets out of balance. Data centre and system architectures adjust to accommodate this.

Compute

Density has multiplied by a factor of 10 in 10 years (slow down of Moore’s Law), made up for this by massive scale in the DC

Networking

Multiplied by 10 in 8 years, 10Gbps about 10 years ago, and 100Gbps about 2 years ago

Data

  • Multiplied by a factor of 1000
  • Storage vendors just haven’t kept up
  • Storage on white boxes?

Pure came in and bought balance to this picture, allowing storage to keep up with networking and compute.

It’s all about

  • Business model
  • Customer experience
  • Technology

“Data is the most important asset that you have”

Pure Storage became a billion dollar company in revenue last year (8 years in, 5 years after it introduced its first product). It’s cashflow positive and “growing like a bat out of hell” with over 4800 customers. Had less than 500 customers just 4 years ago. And a large chunk of customers are cloud providers. Also in 30% of the Fortune 500.

Software

The software economy sits on top of the compute / network / storage stool. Companies are becoming more digital. Last year they talked about Domino’s, and this year they’re using AI to analyse your phone calls. Your calls are being answered by an AI engine that takes your order. An investment bank has more computer engineers and developers than they have investment bankers. Companies need to feed these apps with data. Data is where the money is.

DC Architectures

  • Monolithic scale-up – client / server (1990s)
  • Virtualised – x86 + virtualisation (2000s)
  • Scale-out  – cloud (2010s)

Previously big compute – apps are rigid, now there’s big data – apps are fluid, data is shared

“Data-centric” architecture

  • Faster
  • 100% shared
  • Simpler
  • Built for rapid innovation

Dedicated storage and stateless compute

Examples

  • Large booking, travel, e-commerce site
  • PAIGE.AI – cancer pathology – digitised samples from the last decade
  • Man AHL – economy and stock market modelling

Behind all these companies is a “data hero”

Over 80% of CxOs believe that the speed of analysing data would be one of their biggest competitive issues, but CIOs worried about not being able to keep up with data coming in to the enterprise.

“We empower innovators to build a better word with data”

Beyond AFA

  • Modern data pipeline
  • A whole new model
  • Pure “on-demand”
  • The AI Era

“New meets Now”

It takes great people to make a great company – the amazing “Puritans”. Pure have a NPS score of 83.7 – best in B2B.

 

Matt Kixmoeller

Matt Kixmoeller takes the stage. We need a new architecture to unlock the value of data. Back in 2009. Michael Jackson died, Obama was in, Fusion-IO had just started. Pure came along and had the idea of building an AFA. Today we’re going to bring you the sequel

  • There’s basically SAN / NAS and DAS (which has seen a resurgence in web scale era)
  • DAS reality – many apps, rigid scaling, either too much storage or too much compute

New technologies to re-wire the DC

  • Diverse, fast compute (CPU, GPU, FPGA)
  • Fast networks and protocols (RoCE, NVMe-oF)
  • Diverse SSD
  • Eliminates the outside the box penalty
  • Gets CPUs totally focussed on applications

What if we can finally unite SAN and DAS into a data-centric architecture?

Gartner have identified “Shared accelerated storage”. “The NVMe-oF protocol … will help balance the performance and simplicity of direct-attached storage (DAS) with the scalability and manageability of shared storage”.

“Tier 0”? – they’re making the same mistake again. Pure are focused on shared accelerated storage available for all.

Tomorrow can look like this

  • Diskless, stateless, elastic compute (continuers, VMs, bare metal)
  • Shared accelerated storage (block, file, object)
  • Fast, converged networks
  • Open, full-stack orchestration

 

Keith Martin

Keith Martin (ServiceNow) takes the stage

  • Dealing with high volumes of data
  • Tremendous growth in net new data
  • 18 months ago, doing basic web scale, DAS architecture
  • Filling up DCs at a very fast clip
  • Stopped and analysed everything there was

What happens in an hour in the DC?

In one hour our customers:

  • 7.5 million performance analytics scores computed
  • 730,000 configuration items added
  • 274,000 notifications sent
  • 76,000 assets added
  • 49,200 live feed messages
  • 36,300 change requests
  • 15,600 workflow activities

Every hour of the day our engineering teams:

  • Develop code across the globe in 9 global develoipment locations (SD, SC, SF, Kirkland, London, Amsterdam, Tel Aviv, Hyderabad, Bangalore)
  • Use 450 copies of ServiceNow for quality engineering testing
  • Run 100,000 automated quality engineering tests

In one hour on our infrastructure

  • 25 billion database queries
  • 112 million HTTP requests
  • 2.5 million emails
  • 25.3 million API calls
  • 493TB of backups

We were going through

  • 30K hard drives
  • 3500+ servers
  • >2000 failed HDDs per year

CPU time was being consumed with backup data movement and restore times were becoming longer and longer. They started to look at the FlashBlade. With its small footprint and low power it was a really interesting option for them. It was really easy to setup and use. They let the engineers out of their cages to play with it in the lab and found it was surprisingly hard to break. So they’ve decided to start using FlashBlade in production as their standard for protection data.

Achieving 3x density now

Each rack has:

  • 30 1RU servers
  • 1000 compute cores
  • 1.5PB effective Flash

Decided to test and implement FlashArray as well and they’re excited about FlashArray//X. ServiceNow cares about uptime. Pure has the best non-disruptive upgrade, expansion and repair model. DAS can prove to be expensive at scale.

 

Matt Kixmoeller

Kix takes the stage again

  • 2016: FlashBlade – the world’s first AFA for big data
  • 2017: FlashArray//X

Introducing the FlashArray//X Family

  • //X10
  • //X20
  • //X50
  • //X70
  • //X90

 

Bill Cerreta

Bill Cerreta takes the stage.

  • The FlashArray was launched in 2012, Purity was built to optimise Flash
  • //M chassis designed for NVMe
  • Deep integration of software and hardware

Where are we going with Flash?

SCM, QLC. We’ve eliminated translation layers. The X//90, for example, has

  • Dual-Protocol controllers – speaks to both SSD and NVMe
  • The 10 through 90 have 25GbE onboard
  • Everything’s NVMe/oF ready and this will be added via software later in the year
  • Double the write bandwidth of //M
  • This year, they’re all in on //X
  • 7 generations of evergreen, non-disruptive upgrades [photo]
  • //X makes everything faster (compared to //M)

Neil Vachharajani takes the stage briefly to talk MongoDB on shared accelerated storage.

Kix continues.

Priced for mainstream adoption

  • Early attempts at NVMe cost 10x more than AFAs
  • //X, when introduced last year, was 25% more than //M
  • $0 premium for //X over //M on an effective capacity basis

[Customer video – Berrios]

 

Jason Nadeau

Jason Nadeau takes the stage. Most infrastructure wasn’t built to allow data to flow freely.

  • 10s of products
  • Complex design
  • Silos, difficult to share

“Data-as-a-Service”

Data-centric Architecture

  • Consolidated and simplified
  • Real-time
  • On-demand and self-driving
  • Ready for tomorrow
  • Multi-cloud

Foundation

  • FlashArray
  • FlashBlade
  • FlashStack
  • AIRI

API-first model and software at the heart of the architecture.

 

Sandeep Singh

Sandeep Singh takes the stage. A lot of companies have managed to virtualise. A lot have managed to “flash-ify”. But a lot of them have yet to automate and “service-ize”, to “container-ize”, or to adopt multi-cloud.

Automate and service-ize – on every cloud platform

  • VMware SDDC – VMware SDDC validated design
  • Open automation – pre-built open full-stack automation toolkits
  • Openshift PaaS – container-based reference architecture

Simon Dodsley takes the stage to talk with Sandeep about MongoDB deployments in less than a minute (down from 5 days).

Sandeep continues. Container adoption is increasing quickly but there’s a lack of storage support for persistent containers. Pure have container plug-ins for Docker, Kubernetes. Containerized apps want to consume storage as-a-service. Introducing Pure Service Orchestrator.

Multi-cloud

Introduced ActiveCluster last year. Snapshots and snapshot mobility (portable snapshots introduced last year) are important.

  • Snap to NFS is generally available now
  • CloudSnap to AWS S3 (available in late 2018)
  • DeltaSnap open API (Veeam, Catalogic, actifiio, CommVault, Rubrik, Cohesity)

 

Jason Nadeau

Jason Nadeau comes back on stage. Data as-a-service consumption. Leases aren’t pay per use and aren’t a service-like experience

Introducing Evergreen Storage Service (ES2)

  • Pay per used GB
  • True open
  • Terms as short as 12 months
  • Always evergreen
  • Onboard in days
  • Always “better-than-cloud” economics

Capex with Evergreen storage, Opex with ES2

[Video on PAIGE.AI]

 

Matt Burr

Matt Burr takes the stage. Unlocking the value of what was once cold data. New era demands a new data mindset.

  • How has the value of data changed?
  • How can you extract that value?
  • How can you get started today?

A robot will replace a human surgeon. A machine has learned to adapt faster than the human brain can. More and more data will live in the hotter tier. What tools can make this valuable? Change in the piggy bank – like data. But data is stuck in silos.

  • Data warehouse
  • Data lake
  • Modern data pipeline
  • AI data pipeline

$/GB used to make sense. We need new metrics. $/flops? $/simulation. Real value is generated by simplifying and accelerating the data flow. Build a data hub on FlashBlade. FlashBlade is 16 months old (GA in January 2016).

Invites NVIDIA’s Rob Ober on stage

 

Rob Ober

“The time has come for GPU computing”

  • Moore’s Law is flattening an awful lot
  • NVIDIA as “the AI computing platform”
  • “The more you buy, the more you save”

Traditional hyper scale cluster – 300 dual-CPU servers, 180KW power, or you can deploy 1 DGX-2, 10KW.

Science fiction is being made possible

  • Ultrasound retrofit
  • 5G beam
  • Molecule modelling 1/10 millionth $

Scaling AI

  • Design guesswork
  • Deployment complexity
  • Multiple points of support

AI scaling is hard, “not like your traditional infrastructure”

AIRI

  • Jointly-validated solution
  • Faster, simplified deployment
  • Trusted expertise and support

 

Matt Kixmoeller

Kix takes the stage again. There’s a big gap in AI infrastructure, with customers spread across varying stages of journey from single server -> scale-out infrastructure. Introduces AIRI Mini and they’re also extending AIRI to Cisco.

 

Data Warehouse pitfalls

  • Performance not keeping up with data
  • Pricing extortions and over-provisioning
  • Inflexible appliances built for a single workload

Progress has to have a foundation.

Customer example of telco in Asia moving from Exadata to FlashBlade

Introducing FlashStack for Oracle Data Warehouse

Set your data free

 

Dave Hatfield

Dave Hatfield takes the stage. Thanks for coming. Over 5000 people in the Bill Graham Civic Auditorium and a lot watching on-line. Customers, partners. Be sure to check out the “petting zoo” (Solutions Pavilion). We wanted to have something that was “not your father’s storage show. Your father’s storage show happened last month”. Anyone been to a Grateful Dead show? It’s a community experience, you don’t know what will happen next.

And that’s a wrap.

Storage Field Day Exclusive at Pure//Accelerate 2017 – FlashBlade 2.0

Disclaimer: I recently attended Storage Field Day 13.  My flights, accommodation and other expenses were paid for by Tech Field Day and Pure Storage. There is no requirement for me to blog about any of the content presented and I am not compensated in any way for my time at the event.  Some materials presented were discussed under NDA and don’t form part of my blog posts, but could influence future discussions.

 

These are my rough notes from a session I attended on “Day 0” of Pure//Accelerate 2017 (aka Storage Field Day Exclusive at Pure//Accelerate 2017). Videos of the session can be found here and you can grab my raw notes from here. I try to avoid dumping a bunch of dot points in Tech Field Day posts, but as this one covered some key announcements, I thought most of the information was useful presented as is.

 

A Year of FlashBlade

Par Botes spoke to us briefly about the progress made with FlashBlade in the past year. Originally internally codenamed “Wedding Cake” as it was a white box, the performance is “better than you think”, going from 500K IOPS and 15GB/s, to 1.2M IOPS and 15GB/s, to 1.5M IOPS and 16GB/s in the first six months since GA. I first encountered FlashBlade in the flesh at Storage Field Day 10. You can read more about that here.

 

Scaling Beyond 15 Blades

Rob Lee took some time to talk about scaling beyond 15 blades in a chassis.

  • Linear capacity scale – single namespace growing to dozens of PB scale
  • Linear IOPS and throughput – single namespace / IP scales IOPS and throughout with capacity
  • Preserve simplicity – more capacity adds IOPS & throughout with zero administration

 

Logical View

  • Fabric – scale raw bandwidth without adding management
  • Processing – software dynamically schedules processing resources globally
  • Data – place data as a single system across all blades

 

High level Architecture

  • Integrated Networking – combined internal and external networks. Load balance connections across all blades.
  • Distributed control – partition and distribute control of namespace, data, and metadata across all blades
  • Distributed Data – Distribute persistent data across all blades – high-frequency transactions in NVRAM and longer-lived data in N+2 erasure-coded flash

[image courtesy of Pure Storage]

 

FlashBlade Data Distribution

  • Wide-stripe erasure coding with +2 redundancy
  • Scaling past a single chassis
  • External network load balancer, inter-chassis network switching
  • Added external Fabric Module
  • intra-chassis network switching
  • External Flash Module is 2 rackmounted switches

 

Scaling Fabric Bandwidth

32port 100Gbs switch (1.6Tb/s north-south). Here’s a photo of Rob talking about this.

  • Controller load balancing and capacity load balancing
  • East-west traffic
  • NVRAM/Flash data access
  • Metadata coordination
  • 1.6Tbs across chassis, 300Gbs within chassis

 

Control Placement

Adding a blade is straightforward

  • Partitions rebalance to new blade – stops running on old blade and boots on new blade
  • No data movement required, only compute (data stays in-place)
  • Partition load balancing on a per-blade basis – not chassis constrained

 

Data Placement

  • Data erasure coded across n+2 RAID stripes
  • 15 blades – 13-wide stripe (11+2 parity shards)
  • RAID Groups are dynamic – selected as needed
  • RAID Groups can cross chassis boundaries

As you fill the chassis, it becomes beneficial to constrain the RAID group to a chassis. Note also that there’s enhanced resiliency (n+2 per chassis, without additional overhead) and reduced inter-chassis bandwidth requirements for rebuild operations.

 

Takeaways

  • Software creates parallelism and scale; hardware enables access to data
  • Software/hardware integration without tight coupling
  • Simplicity/reliability created by software control of the network fabric

 

Native Objects

Brian Gold presented a section of the session on Why Objects?

Next Generation Apps

  • Cloud-native development
  • Rich metadata databases

Performance

  • Large & streaming: AI training, media serving, analytics
  • Small & random: time-series metrics, real-time streams

Efficiency

  • No visible partitions
  • Unified management

 

Classic Object Gateways

  • Object API gateway -> file system (index to track metadata)
  • File system becomes bottle neck when scaling to billions of object
  • Purity (FlashArray and FlashBlade) – objects at the core

 

Object Read Path

Request Arrival

  • Extract bucket and object names from request
  • Decode bucket and object names
  • Get bucket ID from authority
  • Get object ID from bucket authority
  • Forward read request to object authority

Read data

  • Read object data from flash
  • Forward back to protocol handler
  • Decompress and form response to client

Two takeaways

  • Two phases – metadata lookup and data access; distributed everything
  • Basically identical to how a file is read via NFS

FlashBlade is S3-compatible for the moment. Purity is really a key-value database

 

Looking Forward

The next generation of applications require new storage interfaces. There was a demo using TensorFlow.

  • Converting raw pixels (ultimate in unstructured data) to structured data
  • Now imagine if you’ve got 10s of thousands of cameras
  • Object detection -> message queue -> object indexing, streaming queries, time-series analysis

 

Conclusion

Par wrapped up by talking about:

  • “The big bang of intelligence”
  • Modern Compute – parallel architecture driving performance
  • New Algorithms – modern approaches for superhuman accuracy
  • Big Data – Data is the new oil
  • “Massively parallel is the new normal”
  • 4th Industrial Revolution (2010 – now) – AI, Big Data, Cloud, IoT, Computing, digital to intelligence

I was a bit confused by FlashBlade when I first heard about it, and suggested that the 12 months post Storage Field Day 10 would be critical to the success of the product. Pure have managed to blow me away with the progress they’ve made with the product since GA, the breadth of customers and use cases they’ve lined up, and the overall level of forward thinking that’s gone into the product. You can use it to do some really cool stuff. The biggest problem I’ve had with the “data is the new oil” paradigm is that, unlike real oil, a lot of companies don’t actually know what to do with their data. FlashBlade is not going to magically fix this for you, but it’s going to give you some pretty compelling infrastructure that solves some of the problem of how to do stuff effectively with massive amounts of data.

Object storage is the new hot, and has been for a little while. Putting together a product like FlashBlade has certainly gotten Pure into a bunch of accounts where they weren’t traditionally successful with FlashArray. It’s also given their more traditional customers a different option for tackling big data problems. Pure strike me as being fiendishly focused on delivering something special with FlashBlade, and certainly don’t appear to be slowing down when adding new features to the platform. There’s been some really cool features added, including support for 17TB blades (almost by accident) and increasing scalability to 75 blades. I’m looking forward to seeing what’s next for FlashBlade. You can read the blog post about the FlashBlade 2.0 announcement here.

 

Storage Field Day Exclusive at Pure//Accelerate 2017 – Purity Update

Disclaimer: I recently attended Storage Field Day 13.  My flights, accommodation and other expenses were paid for by Tech Field Day and Pure Storage. There is no requirement for me to blog about any of the content presented and I am not compensated in any way for my time at the event.  Some materials presented were discussed under NDA and don’t form part of my blog posts, but could influence future discussions.

 

These are my rough notes from a session I attended on “Day 0” of Pure//Accelerate 2017 (aka Storage Field Day Exclusive at Pure//Accelerate 2017). Videos of the session can be found here and you can grab my raw notes from here. I try to avoid dumping a bunch of dot points in Tech Field Day posts, but as this one covered some key announcements, I thought most of the information was useful presented as is.

 

ActiveCluster

Tabriz Holtz and Larry Touchette took us through an overview of ActiveCluster.

 

What do customers really want?

  • Disaster protection
  • Consistent performance
  • Transparent failover
  • Ease of management
  • Subscription to innovation

They don’t want

  • Life to be difficult because they’re running synchronous replication
  • To pay for it

 

Multi-site Active/Active

  • Zero RPO
  • Zero RTO
  • Zero $
  • Zero Additional Hardware

 

Basic Architecture

  • Symmetric Active/Active
  • Pure’s pod management model – container where you can store your volumes
  • Passive Pure1 Cloud Mediator – prevent split brain

 

Pods

  • Simple management model
  • only 1 new command introduced
  • serves as a container and a consistency group
  • keeps metadata with its data

 

4 Steps to Setup

The whole point is that it’s super simple to setup, so much so that you can do it in four steps from the CLI.

1. Connect the arrays

purearray connect --type sync-replication

2. Create a stretched pod

purepod create pod1
purepod add --array arrayB pod1

3. Create a volume

purevol create --size 1T pod1::vol1

4. Connect your hosts

purehost create --preferred-array arrayA host
pure host connect --vol pod1::vol1 host

 

Symmetric Active/Active

I/Os perform symmetrically

  • 1 round trip for writes
  • reads serviced locally

Host ALUA preferences:

  • Active/Optimised
  • Active/Non-optimised

There’s a 5ms RTT limit and it uses TCP/IP between arrays (Ethernet). Independent dedupe runs on both sides.

 

Passive Mediator

  • No split brain … ever!
  • Intelligence is in the arrays
  • Mediator imply records failover
  • No third site needed
  • Arrays alert if they can’t access mediator

There is also the option to deploy a VM that can be used on-premises. While the cloud mediator runs multiple instances behind a load balancer, the on-premises mediator would have to be protected with HA or similar.

So what if I lose comms to the outside world? (both to the outside world and the partner array). Volumes will be taken offline. The mediator is a per pod setting, so you could conceivably use both in your environment.

 

Transparent Recovery

1. Snapshots sent asynchronously until arrays are nearly in sync

2. IOs forwarded synchronously along with final snap are merged into target

3. Arrays are fully in sync with no pause in IO for final sync

The goal is to allow different arrays to replicate with each other. At GA these will be qualified. Purity versions (1 minor version different e.g. 5.1 and 5.2). Every customer gets this feature (yay, Evergreen). Assuming you have the appropriate supporting infrastructure. And there’s support for multiple connection types.

 

DirectFlash Shelf

Pete Kirkpatrick (Chief Hardware Architect) spoke about the recently announced FlashArray//X.

  • 100% NVMe Enterprise AFA
  • “This is just a flash array”
  • The data is the array, the hardware and software comes and goes over time

They were working on the FlashArray//M chassis about 4 years ago, and had gone to some length to future proof the design. “DirectFlash Modules” are now replacing the SAS SSDs. SSDs emulate HDDs – this isn’t ideal.

 

Flash Transition Layer needs Garbage Collection

  • Severely limits sustained throughout
  • Destroys latency distribution
  • Causes excess wear
  • Needs over provisioned capacity

 

Legacy protocols and interfaces

  • Assumed high latency
  • Inherently serialised

 

DirectFlash

Purity has always been designed for Flash, so get disk legacy out of the way. The goal of DirectFlash was to

  • Start with NVMe: efficient, low latency, high throughput, high parallelism,
  • Design an API providing knowledge of the Flash geometry
  • Place data intelligently, and schedule operations with high precision
  • No FTL is required, so GC is eliminated
  • High sustained throughput and low, deterministic latency
  • No over provisioning
  • Extended flash endurance

Here’s a happy snap of one of the 18.3TB (I think) modules.

You can now fit 1PB of useable Flash in 3RU. DirectFlash Shelf is this week’s news. They’re using NVMe/F. It’s over RoCE (RDMA over Converged Ethernet). You can start with 1 shelf at this stage, but Pure are looking to extend that capability.

 

VVols Support

Cody Hosterman took us through VMware Virtual Volumes (VVols) support with Purity//FA 5.0. I enjoy watching Cody present and I wasn’t disappointed by this session. So, VVols eh? Why?

  • Virtual disk granularity on array – use array-based technology on a virtual disk basis
  • Automatic volume creation and configuration
  • Storage Policy Based Provisioning

 

Virtual Volumes – The Full Picture

[image via Pure Storage]

 

VVols

Every VM has individual volumes on the array:

  • Config VVol—4 GB—holds the configuration information of the VM. Created automatically when a VM is provisioned
  • Swap VVol—is for the VM swap file. Sized according to the VM memory. Created automatically when the VM is powered-on and deleted when powered-off
  • Data VVol—for every virtual disk added to the virtual machine there is a new data VVol. Sized by the requested size of the virtual disk

 

VVol Snapshots

  • VMware snapshot and array-snapshot is created automatically – no performance penalty

 

The Data Plane

Protocol Endpoints

  • A mount-point for VVols
  • Presented in a traditional fashion via iSCSI or FC as LUN
  • VVols are sub-LUNs to a PE. IO goes to the PE on the array, the array distributes.
  • FlashArray automatically “binds” VVols to the appropriate PE

 

The Management Plane

  • How does vCenter manage the FlashArray? Through a VASA provider

FlashArray VASA Provider

  • VASA version 3 (includes replication)
  • Redundant service on both controllers
  • Automatically configured during Purity upgrade
  • Active-active configuration
  • Entirely stateless – no configuration is tied to /stored on the hardware of the controllers (no special VASA database on the array, it is part of the array configuration)

 

Policy Based Management

Create VMs and individual virtual disks (VVols) independently

Use customised capabilities advertised by VASA provider to configure volumes

  • Replication
  • Snapshot policy
  • QoS
  • Etc.

Are VVols Special Volumes? On the FlashArray, not really. Just normal volumes with special metadata tags. So if you want to present a VVol to a physical server for instance? You can just connect it as a standard LUN.

 

Conclusion

I’m enthusiastic as all get out about ActiveCluster. There have been rumblings in the market about this type of capability for some time, so it’s great to see Pure deliver. I need to dig a bit deeper into it, but it feels a lot like it has the cross-site capability of Dell EMC’s VPLEX without a lot of the palaver traditionally associated with that product (which to be fair, does more than just cross-site volumes). The great thing is that it’s available to existing customers without additional expense (at least on the Pure side) or messing about. This has always been a big selling point for Pure, and it’s great to see it continue here.

I think the DirectFlash shelf is certainly a step in the right direction and the appetite is there for this kind of solution. I’ll be interested to see how many shelves they end up adding, as the scale and speed possibilities here are potentially pretty tremendous. It will also be interesting to see the uptake of the solution over the next 12 months.

I liked a lot of what I saw with Cody’s presentation on VVols support. It certainly appeared fairly straightforward. I remain underwhelmed by VVols in general though. I know there needed to be a change in the way we presented storage to VMs but it feels like we’ve somehow missed the boat with this solution. In another year it might just be that everything sits on VVols by default but I feel like that’s been the feeling for the past five years and it hasn’t yet transformed to the extent we expected. I am more than happy to be proven wrong on this point though, and my surliness regarding VVols shouldn’t be taken as criticism of what Pure have managed to deliver here. Also, it’s worth checking out Cody’s post on Virtual Volumes support here – he covers it way better than I do.

Storage Field Day Exclusive at Pure//Accelerate 2017 – General Session Notes

Disclaimer: I recently attended Storage Field Day 13.  My flights, accommodation and other expenses were paid for by Tech Field Day and Pure Storage. There is no requirement for me to blog about any of the content presented and I am not compensated in any way for my time at the event.  Some materials presented were discussed under NDA and don’t form part of my blog posts, but could influence future discussions.

 

Here are my General Session notes from Day 1 of Pure//Accelerate 2017. I’ll be digging into some of the announcements in the very near future so this is just a rough teaser. I call it death by dot point.

 

David Hatfield

David Hatfield in a Dubs cap. Sorry about the traffic, but it can be hard “[w]hen you setup an event for 3000 people in kind of a crack area?.” He then talks about Game 5 of the NBA Finals. Thanks to people for coming a long way (some travelled around 36 hours). 227800 square feet of space. Used for manufacture of iron originally. This will be the last event ever in this building. Will be knocked down soon. There’s also a 110 feet wide screen, 420 million pixels, FlashBlade helps to render.

“New and Possible”

  • 25 new software capabilities, new hardware, new cloud capabilities, new partners (friendships?)
  • “The new possible” – Help you break free from legacy way of doing things.
  • New – look around and identify what real innovation looks like.

Partner and sponsor shoutout

  • It takes an ecosystem of partners to help too, and sponsors as well.
  • Shoutout to Veeam – will be directly integrating Veeam and Pure Storage in next release of Availability Suite (v10?)
  • Cisco – FlashStack (over 1400 customers together). 7 CVDs in place at the moment. 2 new offers – NVMe over Fabric to the host, Cisco Capital – full offering for FlashStack

 

Scott Dietzen

Scott Dietzen (CEO) takes the stage, and reclaims his Dubs cap. Am I in the right place? I knew I was when I got inside. Blue is the colour of so many of our competitors, but it’s a Warriors cap, so it’s okay.

This is our second //Accelerate conference. Thanks to 3300 customers (and partners).

“The world’s most valuable resource is data”

  • Companies are competing to amass huge datasets (are they doing useful things with it though?)
  • AI rates only behind cloud and mobile in terms of impact people think it will have
  • “Massively parallel AI demands massively parallel storage”

“By the year 2020 the amount of data created will be 50+ zettabytes”

  • Capacity of the Internet will only be 2.5 zettabytes
  • It’s going to stored close to where it’s generated

A new model for DCs

  • Multi-cloud
  • Core
  • Edge

Pure uses tonnes of SaaS to run its business, but it also has its own DCs. Believes Edge is going to be larger than multi-cloud and core together because of IoT, etc.

 

Disruptors? (based on a recent survey)

  • The digital gold rush
  • The great workload debate
  • Cloud: migrations and mistakes
  • Data: taking back control

Multi-cloud = photo of Snoop Dogg blazing. Pure are looking to deliver “The data platform for the cloud era”.

 

Differentiators?

  • Big data bandwidth
  • Performance for deep learning
  • ultra-density
  • Uptime
  • Subscription to innovation

 

Tomorrow’s cloud block:

  • 1300 cores
  • 2.6PB Flash
  • 2x density and 5x performance improvement (for leading SaaS company)
  • 100x reduction in space for enterprise DC (20 racks -> 4RU)

$1B revenue and cash flow positive, 6th year in selling

All the new software updates are given to customers for free

 

Liz Centoni

Liz Centoni (Cisco SVP and GM, Computing Systems Product Group) takes the stage to talk with Scott Dietzen.

Every DC was built to do one thing: run applications. But these applications are changing – how they’re built, where they reside, etc. They’re a lot more distributed. A lot more endpoints to manage and secure.

How does FlashStack help?

  • Helps customers build their private cloud infrastructure
  • This year – adding hybrid cloud (via Cisco CloudCenter)

FlightStats example

Servers existed before DCs or cloud. Customers want any workload, anywhere

  • Compute is moving closer to the data.
  • Security is top of mind for everyone.
  • UCS is a “system” – fabric-centric design, 100% programmable, 60K customers.
  • Start small and scale

What about the impact of NVMe?

  • SSDs changed the storage bottleneck, but NVMe really puts it back on the network. Cisco is happy about having the opportunity to improve the performance of the whole stack.

 

Kelli Zielinski

Kelli Zielinski, Domino’s takes the stage. Traditional arrays just weren’t keeping up. Invested in Pure Storage FlashStack. Now “[t]he application gets what it needs”.

 

Matt Kixmoeller

Matt Kixmoeller takes the stage. It’s the “dawn of a new cloud era – yet the old never really disappears”.

Purity

  • Reduce
  • Assure
  • Protect
  • Secure
  • Direct Flash

Pure1

  • Manage
  • Analyse
  • Support
  • Meta

FlashBlade for big data, FlashArray//X for low latency/high IOPS apps

It’s time for software to take centre stage – 25 new features (delivered in an Evergreen fashion)

 

Tier 1

  • “no one ever got fired for buying [blue]”.
  • You want reliability and innovation (dedupe, compression, simple, automated, open cloud integration, NVMe, NVMe/F)
  • Delivered 2 years of 6 9s since GA

Metro Stretch Cluster – 1994 – EMC launched SRDF – 20 years later – still hard and expensive.

 

Purity//FA 5.0

Purity ActiveCluster.

Steve Hodgson (Software Architect) does a brief demo on setup.

 

Jason Nadeau

Jason Nadeau takes the stage.

  • Compression 2.0 – 25% improvement – self-selecting compression engines
  • Simplest VVols implementation in the industry
  • Granular VM-level operations and transparency
  • Cloud automation and security compliance

 

Purity//FA Snap and CloudSnap

“Portable snapshot”

  • Snap Local
  • Snap to FlashArray
  • Snap to FlashBlade
  • Snap to NFS
  • CloudSnap to AWS
  • DeltaSnap to API

So now you can:

  • Native, two-way cloud connection
  • Backup, restore, migrate, and DR
  • Fully leverage all PaaS services

 

Purity //Run

  • Run VMs and Containers Directly on Purity
  • Ideal for Edge Analytics, Custom protocols
  • Flexible, open, secure, HA platform

 

Windows File Services for Purity

  • Best of Breed – FlashArray meets MS File services

 

DirectFlash Shelf

  • Native NVMe/F expansion shelf (photo)

 

“The new Tier 1 is Evergreen”

 

Arthur Riel

Arthur Riel (Director, The World Bank) takes the stage. The World Bank is neither Wall Street nor Main Street. Their job is to decrease poverty. When he got there, they were a risk-averse organisation (“Let’s keep doing what we’re doing because it works”).

  • But “[i]f I save money and can’t deliver my services what good is that?”.
  • “Slow storage covers up a lot of sins up the stack”

 

Sumit Dhawan

Sumit Dhawan (runs EUC at VMware) talks with Scott Dietzen.

“Tech is going out of tech” you need a more platform-centric approach (?)

Dietzen: Is it hard to work with us when your parent company is Dell? (I’m paraphrasing).

 

Par Botes and Rob Lee

Par Botes and Rob Lee take the stage to talk about FlashBlade.

  • “The big bang of intelligence”
  • Medium blade – 17TB (fits in between 8TB and 52TB)
  • 75 blade-scale FlashBlade (start as small as 7 blades, scale 1 at a time)

 

Why Object Storage?

  • Cloud-native applications use object storage
  • Next-generation developers code with object
  • Cloud primary storage is object
  • >10x faster time to first byte vs S3
  • >100x faster indexing image objects vs existing solution at leading web scale company

Brian Gold joins them on stage – it’s about AI from edge to cloud

 

Rob Ober

Rob Ober (Tesla Chief Platform Architect, Nvidia) takes the stage with Scott Dietzen.

Deep learning (about 5 years ago) has taken off because of:

  • Neural nets (this have been around a while)
  • Massive amounts of data (tremendous volumes)
  • Computation

“You need good data”

 

Sandeep Singh

Sandeep Singh on stage to talk about self-driving storage

  • Automate and simplify
  • Sense and model world around
  • Constantly learn and re-train
  • Global effect

 

>7PB telemetry data

> 1 trillion data points per day

 

Pure1

>500 Sev1 incidents avoided to date

 

Performance sizing has been the final frontier

  • Too many variables
  • Complex interaction and inter-dependencies
  • Over provisioning = wasted expense
  • under-provisioning = downtime

This is a perfect problem for AI and machine learning.

 

Pure1 Meta

  • Global sensor network
  • real-time scanning
  • data lake
  • AI engine

 

Sergey Zhuravlev, Chief Data Scientist

  • >1000 measures
  • “workload DNA”
  • Meta learns from everyone’s workloads to make better predictions
  • Meta Workload Planner

 

David Hatfield takes the stage to wrap up. Here’s a summary of today’s announcements.

 

Purity//FA 5.0

  • ActiveCluster
  • Snap to FlashBlade
  • Windows File Services for FlashArray (SMB and NFS)
  • Policy QoS
  • Snap to NFS
  • Compression 2.0
  • VVols
  • Hybrid Cloud for AWS
  • Microsoft ODX
  • CloudSnap to AWS
  • Purity /Run
  • Docker Persistent Volumes

 

Purity//FB 2.0

  • Object / S3
  • SMB
  • Snapshots
  • Scale to 75 Blades
  • HTTP
  • REST
  • LDAP
  • IPv6
  • NFS NLM

 

Pure1

  • Pure1 Meta
  • Workload Planner
  • Cloud Mediator
  • Global Dashboard
  • Workload DNA
  • Cloud REST

 

Good session. 4 stars. Stephen did a nice live blog as well – you can read it here.

 

Storage Field Day – I’ll Be At Storage Field Day 13 and Pure Accelerate

Storage Field Day 13

In what can only be considered excellent news, I’ll be heading to the US in early June for another Storage Field Day event. If you haven’t heard of the very excellent Tech Field Day events, you should check them out. I’m looking forward to time travel and spending time with some really smart people for a few days. It’s also worth checking back on the Storage Field Day 13 website during the event (June 14 – 16) as there’ll be video streaming and updated links to additional content. You can also see the list of delegates and event-related articles that have been published.

I think it’s a great line-up of presenting companies this time around. There are a few I’m very familiar with and some I’ve not seen in action before.

*Update – NetApp have now taken the place of Seagate. I’ll update the schedule when I know more.

 

I won’t do the delegate rundown, but having met a number of these people I can assure the videos will be worth watching.

Here’s the rough schedule (all times are ‘Merican Pacific and may change).

Wednesday, Jun 14 09:30-10:30 StorageCraft Presents at Storage Field Day 13
Wednesday, Jun 14 16:00-17:30 NetApp Presents at Storage Field Day 13
Thursday, Jun 15 08:00-12:00 Dell EMC Presents at Storage Field Day 13
Thursday, Jun 15 13:00-14:00 SNIA Presents at Storage Field Day 13
Thursday, Jun 15 15:00-17:00 Primary Data Presents at Storage Field Day 13
Friday, Jun 16 10:30-12:30 X-IO Technologies Presents at Storage Field Day 13

 

Storage Field Day Exclusive at Pure Accelerate 2017

You may have also noticed that I’ll be participating in the Storage Field Day Exclusive at Pure Accelerate 2017. This will be running from June 12 – 14 in San Francisco and promises to be a whole lot of fun. Check the landing page here for more details of the event and delegates in attendance.

I’d like to publicly thank in advance the nice folks from Tech Field Day who’ve seen fit to have me back, as well as Pure Storage for having me along to their event as well. Also big thanks to the companies presenting. It’s going to be a lot of fun. Seriously.