Bringing confidential computing to Kubernetes

Historically, data has been protected at rest through encryption in data stores, and in transit using network technologies, however as soon as that data is processed in the CPU of a computer it is decrypted and in plain text. New confidential computing technologies are game changing as they provide data protection, even when the code is running on the CPU, with secure hardware enclaves. Today, we are announcing that we are bringing confidential computing to Kubernetes workloads.

Confidential computing with Azure

Azure is the first major cloud platform to support confidential computing building on Intel® Software Guard Extensions (Intel SGX). Last year, we announced the preview of the DC-series of virtual machines that run on Intel® Xeon® processors and are confidential computing ready.

This confidential computing capability also provides an additional layer of protection even from potentially malicious insiders at a cloud provider, reduces the chances of data leaks and may help address some regulatory compliance needs.

Confidential computing enables several previously not possible use-cases. Customers in regulated industries can now collaborate together using sensitive partner or customers data to detect fraud scenarios without giving the other party visibility into that data. In another example customers can perform mission critical payment processing in secure enclaves.

How it works for Kubernetes

With confidential computing for Kubernetes, customers can now get this additional layer of data protection for their Kubernetes workloads with the code running on the CPU with secure hardware enclaves. Use the open enclave SDK for confidential computing in code. Create a Kubernetes cluster on hardware that supports Intel SGX, such as the DC-series virtual machines running Ubuntu 16.04 or Ubuntu 18.04 and install the confidential computing device plugin into those virtual machines. The device plugin (running as a DaemonSet) surfaces the usage of the Encrypted Page Cache (EPC) RAM as a schedulable resource for Kubernetes. Kubernetes users can then schedule pods and containers that use the Open Enclave SDK onto hardware which supports Trusted Execution Environments (TEE).

The following pod specification demonstrates how you would schedule a pod to have access to a TEE by defining a limit on the specific EPC memory that is advertised to the Kubernetes scheduler by the device plugin available in preview.

Now the pods in these clusters can run containers using secure enclaves and take advantage of confidential computing. There is no additional fee for running Kubernetes containers on top of the base DC-series cost.

The Open Enclave SDK was recently open sourced by Microsoft and made available to the Confidential Computing Consortium, under the Linux Foundation for standardization to create a single uniform API to use with a variety of hardware components and software runtimes across the industry landscape.

Try out confidential computing for Kubernetes with Azure today. Let us know what you think in our survey or on GitHub.
Quelle: Azure

Finastra “did not expect the low RPO” of Azure Site Recovery DR

Today’s question and answer style post comes after I had the chance to sit down with Bryan Heymann, Head of Cloud Architecture at Finastra, discussing his experience with Azure Site Recovery. Finastra builds and deploys technology on its open software architecture, our conversation focused on the organization’s journey to replace several disparate disaster recovery (DR) technologies with Azure Site Recovery. To learn more about achieving resilience in Azure, refer to this whitepaper.

 

 

You have been on Azure for a few years now – before we get too deep in DR, can you start with some context on the cloud transformation that you are going through at Finastra?

We think of our cloud journey across three horizons. Currently, we're at “Horizon 0” – consolidating and migrating our core data centers to the cloud with a focus on embracing the latest technologies and reducing total cost of ownership (TCO.) The workloads are a combination of production sites and internal employee sites.

Initially, we went through a 6-month review with a third party to identify our datacenter strategy, and decided to select a public cloud. Ultimately, we realized that Microsoft would be a solid partner to help us on our journey. We moved some of our solutions to the cloud and our footprint has organically grown from there.

All this is the enabler to go to future “horizons” to ensure we continuously keep pace with the latest technology. Most importantly we’re looking to move up the value chain – so instead of us worrying about standing a server up, patching a server, auditing a server, identity on a server… we’re now ingesting and deploying the right policies for the service (not the server) and taking advantage of the availability, security, and disaster recovery options.

Exciting to hear about the journey you've taken so far. I believe DR is a requirement across all of those horizons, right?

Disaster recovery is front and center for us. We work closely with our clients to regularly test. At this point we have executed more than 50 test failovers. Disaster recovery and backups are non-negotiable standards in our shared environment.

What were you most skeptical about when it came to DR in Azure? What was it that helped you become convinced that Azure Site Recovery was the right choice?

We used just about every tool in our data centers and always had mixed results. We thought that Azure Site Recovery might be the same, but I was glad that we were wrong. We have a strong success rate and even wrote special dashboards to track our recovery point objective (RPO) for a holistic view on our posture! We were skeptical that Site Recovery would be point and click capable, and whether it would be able to keep up with the amount of change we have, when failing over from the East coast to the West coast. Our first DR test in Azure, over two years ago now, was actually wildly successful. We did not expect the low RPO that we saw and were delighted. I think this speaks volumes to Azure’s network backbone and how you handle replication, to be that performant.

We hear that from a lot of customers. Great to get further validation from you! Could you ‘double click’ on your onboarding experience with Azure Site Recovery, up to that first DR drill?

There wasn't any heavy onboarding, which is a good thing as it really wasn’t needed. It was so intuitive and easy for our team to use. The documentation was very accurate. The point and click capabilities of Site Recovery and the documentation enabled us to onboard and go. It has all been in line with what we needed, without surprises.

What kind of workloads are you protecting with Azure Site Recovery?

All of our virtual machines (VMs) across North America are using Site Recovery, everything from our lending space, to our payment space, to our core space. These systems support thousands of our customers, and each of those have their own end customers which would number in the millions – Site Recovery is our default disaster recovery mechanism across the whole fleet.

Wow, that’s a lot of customers and some sensitive financial spaces so no wonder disaster recovery is such a high priority for your teams. We regularly hear prospective customers asking whether Azure Site Recovery supports Linux – I'd love to understand if you have Linux-based applications using Site Recovery, and what your experience has been with those?

Actually, it was our very first application for which we deployed Azure Site Recovery – and it’s all Linux. Linux support for Site Recovery has been fantastic. We failover every six months, without any issues. The ease of use and the amount of times we have tested now has significantly increased. We pressed on our normal RPOs to get them down to very, very aggressive levels. Some of our Linux based applications are complex, but Site Recovery has worked without any issues.

You touched upon DR drills – I'd love to understand what your drill experience has been like?

The experience has been seamless and simple. The application itself may have some configurations that need to be considered during DR drills, such as post-scripts, but those are hammered out quickly. We try to do drills every six months, but at least once every year.

Which features of Azure Site Recovery do you like the most?

I love that I can fail across regions. I also love the granular recovery point reporting. It allows us to see where we may or may not be seeing problems. I'm not sure we ever got that from any other tools, it’s very powerful and it’s graphical user interface based – and any Joe could do it, it's not hard to select a VM and replicate it to another region. I especially like the fact that we are only charged for storage in the far side region so, financially, there's not an impact of having warm standbys and still we are able to hit our RPO.

If you were to go through this journey all over again with Azure Site Recovery, is there anything that you would have done differently?

I would have liked to get our knowledge base and plans in place for a month longer before implementing it. It's just so easy that we were able to blow through most of it, but we did miss a couple of simple things early on which were easily fixed later on our journey. We found out quickly we didn't want standard hard drives, we wanted premium for example.

Looking forward, how do you plan to leverage Azure Site Recovery?

We recently used Azure Site Recovery to move a customer in our payment space from on premises to Azure – we will now get those machines on Site Recovery across Azure regions, we're not going to rebuild the entire platform. It's obviously the de facto to get us out, and it is the standard for regional disaster recovery for VMs there. There is no other product used.

People ask me what keeps me up at night, there are really two things. “Are we secure?” and “Can we recover?” – I call it the butterfly effect. When you come in each morning, are you confident that if you cratered a datacenter, you could come up in a different one? I can confidently answer that with yes. We could fail out to another region, with all our data. That's a pretty nice spot to be in, especially when you're sitting in a hyperscale cloud. I know that I have storage replication. I know that I own the network links. To allow somebody to run this stuff on our behalf was a mindset change, but it has really been a positive experience for us.
Quelle: Azure

Elastic Fabric Adapter jetzt mit Intel® MPI Library kompatibel

Elastic Fabric Adapter (EFA) ist jetzt mit dem Update 6 2019 der Intel® MPI Library kompatibel. Die Intel® MPI Library ist eine Message Passing Library, die den MPI-Standard (Message Passing Interface) implementiert. Kunden können anhand der Bibliothek Anwendungen erstellen, warten und testen, die besser auf HPC-Clustern mit Prozessoren von Intel® laufen.
Quelle: aws.amazon.com

Verbesserte Suchfunktion für Parameter Store

Ab heute ist in AWS Systems Manager Parameter Store eine verbesserte Suchfunktion verfügbar, mit der sich Parameter einfach nach ihrem Namen suchen lassen. Die neue Suchfunktion vereinfacht das Durchsuchen großer Parametermengen in Ihrem Konto und das Auffinden von Parametern, deren genauen Namen Sie nicht kennen.  
Quelle: aws.amazon.com

Achieve peace of mind with BigQuery pricing and control

For many companies, data analytics has evolved from an occasional task to something that’s mission-critical to their business. When you’re doing data analytics at scale, predictable spending is key—something we often hear from our enterprise customers, like HSBC and Sky.To that end, we launched BigQuery’s flat-rate pricing a few years ago, a fixed-rate pricing model that makes it easy for you to predict and control your monthly BigQuery bill. We’re happy to announce BigQuery Reservations, an easy and flexible self-service way to take advantage of BigQuery flat-rate pricing, available in beta in coming days. Reservations makes it even simpler to plan your spending and add flexibility and visibility to your data analytics use cases. You’ll see this feature in cloud console in the next two weeks.BigQuery Reservations lets you:Control your transparent, predictable BigQuery analysis spending. Purchase BigQuery slots in BigQuery’s web UI in seconds. Seamlessly manage your enterprise workloads in BigQuery.Avoid compute silos by easily sharing idle capacity across your entire organization.BigQuery Reservations solves enterprise customers’ largest problems”BigQuery Reservations will bring a new kind of flexibility and predictability to enterprises doing large-scale data analytics. For a serverless, cloud-native data warehouse like BigQuery, the ability to predict costs for customers is huge. And with our research showing that 42% of organizations plan to use or are exploring serverless analytics over the next 12 months, pricing and consumption flexibility will serve as a key differentiator for GCP,” says Mike Leone, Senior Analyst at Enterprise Strategy Group. “This announcement opens up more possibilities for workload management, and adds higher levels of efficiency with idle slot capacity being available for reuse.”We’ve heard that you need more power and flexibility with your resources, and the ability to buy and manage BigQuery slots on your own. “Reservations were instrumental in helping us incrementally ramp up slot capacity as we migrated over from another data warehouse, greatly increasing our cost performance,” says Jingsong Wang, engineering manager, Discord. “The ability to share idle slot capacity across projects, workloads and users helps ensure our business-critical workstreams stay online, while giving users the flexibility to run more complex workloads.”You can get a full demo here on how BigQuery Reservations work:Reservations can bring solutions to common issues:Cost predictability and conformism to budgets. While cloud-native services offer unparalleled scalability and efficiency, it’s often at the expense of cost predictability, resulting in budget overruns. Pay-per-use pricing models are especially hard to manage. BigQuery Reservations offers a predictable flat-rate pricing model—no surprises on your monthly bill.Immediate access to capacity. With BigQuery Reservations, purchasing slot commitments merely takes seconds. There is no need to wait for your data warehouse to spin up, and you no longer need to warm up your data warehouse’s disk-driven adaptive cache to get optimal performance. Enterprise-grade workload management. Your data science group may run high-priority, high-demand workloads, and need to have isolated and guaranteed analytics capacity, whereas your test workloads need access to only a small amount of capacity. Reservations lets you dynamically and programmatically partition your BigQuery slots into pockets of resources dedicated to departments or workloads.Efficiency. Partitioning analytics capacity can create compute silos, in which capacity is wasted. BigQuery Reservations distribute any unused BigQuery slots in real time to workloads with high-capacity demands, so even the largest and most complex environments can take advantage of every single BigQuery slot at any time. It’s time to say no to compute silos!We’ve heard from media company Sky that they’ve found this pricing useful. “Sky has been using BigQuery’s flat-rate for some time now,” says Vince Marco, enterprise infrastructure architecture manager at Sky. “Taking advantage of BigQuery’s flat-rate pricing has given Sky peace of mind when it comes to performance and our BigQuery bill. Reservations helped Sky rethink how to protect business-critical workloads, while isolating lower-priority development projects and making sure we get the most of BigQuery’s performance.”Adding cost predictability to your data warehouseBigQuery Reservations is a flexible platform for administering resources and workloads. It involves a three-step process to manage your environment:Commitments, which let you purchase slot capacity.Reservations (optional), which give you the ability to partition your capacity.Assignments, which gives you the ability to assign projects, folders, or your entire organization to Reservations.As an example, you may need 1,000 BigQuery slots for your organization. Your BigQuery users include a data science team, a high-priority ELT workload, and BI dashboards. With BigQuery Reservations, you can:Purchase a 1,000-slot commitment Create reservations “ds” with 500 slots, “elt” with 300 slots, and “bi” with 200 slotsAssign the data science team’s Google Cloud projects to “ds” reservationAssign your ELT projects to “elt” reservationAssign the project that runs your BI dashboard to “bi” reservationNow each of your workloads has dedicated capacity. In addition, any single unused BigQuery slot is automatically and immediately available to other workloads in your organization.You can perform these actions right in the BigQuery UI, or programmatically in the BigQuery command-line tool. We’ve heard from customers that BigQuery Reservations can streamline workload management and add efficiency. “The Slot Reservation API strikes a good balance between control and flexibility for managing BigQuery workloads,” says Henry Lin, engineering manager, Reddit. “We’re able to isolate expensive queries from each other without fearing that we’re underutilizing BigQuery resources. The API has been remarkably easy to use and, in turn, has empowered us to optimize our workflows without needing to micromanage them.”Getting started with BigQuery ReservationsBigQuery’s flat-rate pricing starts at 500 slots, and is generally a good fit for production usage and when customers are looking for price predictability. Customers can still take advantage of on-demand, serverless pricing for proof-of-concepts (POCs) and ad-hoc analysis. Here’s a look at when you might use one or the other:To get started with Reservations, head over to the Reservations getting started documentation. What’s nextBigQuery is a serverless enterprise data warehouse. As such, we strive to reduce our users’ administrative overhead, and to automate the day-to-day toil associated with managing a typical data warehouse. Reservations continues that trend by introducing powerful features that give you more control over your BigQuery environment. We are looking forward to hearing your feedback!We’re putting the final touches on BigQuery Reservations. Check back in soon.Learn more about:BigQuery flat-rate pricing documentationReservations Quickstart guideReservations documentationWhat is a BigQuery slot? documentationChoosing between on-demand and flat-rate pricing modelsEstimating the number of slots to purchaseGuide to workload management with Reservations
Quelle: Google Cloud Platform