Better together: Expanding the Confidential Computing ecosystem

Core to our goal of delivering security innovation is the ability to offer powerful features as part of our cloud infrastructure that are easy for customers to implement and use. Confidential computing can provide a flexible, isolated, hardware-based trusted execution environment, allowing adopters to protect their data and sensitive code against malicious access and memory snooping while data is in use. Today, we are happy to announce that we have completed the rollout of Confidential VMs to general availability in nine regions. Our partners have played a huge part in this journey. They have been critical in establishing an ecosystem that aims to make Confidential Computing ubiquitous across mobile, edge, and cloud. We spoke to Raghu Nambiar from AMD, Mark Shuttleworth from Canonical, Burzin Patel from HashiCorp, Mike Bursell from Red Hat, Dr. Thomas Di Giacomo from SUSE, and Solomon Cates from Thales. Here are the excerpts.Raghu Nambiar, Corporate Vice President, Data Center Ecosystems, AMDConfidential Computing is a relatively new concept with a goal to encrypt data in use in the main memory of the system, while still offering high performance. Confidential Computing addresses key security concerns many organizations have today in migrating their sensitive applications to the cloud and safeguarding their most valuable information while in-use by their applications. It wouldn’t be a surprise if in a few years, all virtual machines (VMs) in the cloud are Confidential VMs. How did you approach confidential computing?The 2nd Gen AMD EPYC processors used by Google for its Confidential VMs uses an advanced security feature called Secure Encrypted Virtualization (SEV). SEV is available on all AMD EPYC processors and, when enabled by an OEM or cloud provider, it encrypts the data-in-use on a virtual machine, helping to keep it isolated from other guests, the hypervisor and even the system administrators. The SEV feature works by providing each virtual machine with an encryption key that isolates guests and the hypervisor from one another, and these keys are created, distributed, and managed by the AMD Secure Processor. The benefit of SEV is that customers don’t have to re-write or re-compile applications to access these security features With SEV-enabled Confidential VMs, customers have better control of their data, enabling them to better secure their workloads and collaborate in the cloud with confidence. What kind of performance can we expect?What’s really impressive about Google Confidential VMs powered by AMD EPYC processors with SEV enabled is that it offers performance close to that of non-confidential VMs. AMD and Google’s engineering teams ran a set of well-known application benchmarks for relational database, graph database, webserver as well as Computational Fluid Dynamics and popular simulation workloads in FSI on Google Confidential VMs and Google’s N2D VMs, of which Confidential VMs are based on. The difference in using SEV versus not using SEV on the applications listed, was measured to be just a small overhead in application performance. Any final thoughts?Confidential Computing is a game-changer for computing in the public cloud as it addresses important security concerns many organizations have about migrating their sensitive applications to the cloud. Google Confidential VMs, with AMD EPYC processors and SEV, strengthen VM isolation and data-in-use protection helping customers safeguard their most valuable information while in-use by applications in the public cloud. This is a paradigm shift and we’re excited to work with Google to make this possible.Mark Shuttleworth, CEO, CanonicalConfidential Computing directly addresses the question of trust between cloud providers and their customers, with guarantees of data security for guest machines enforced by the underlying hardware of the cloud. With Google’s addition of Confidential Computing to multiple regions, customers gain a secure substrate for large-scale computation with sensitive data and a path to regulatory compliance for new classes of workload on the cloud.What value does the partnership between GCP and Canonical create?Close technical collaboration between Google and Canonical ensures that Ubuntu is optimized for GCP operations at scale. Confidential Computing requires multiple pieces to align and we are delighted to offer full Ubuntu support for this crucial capability at the outset with Google.How will this benefit organizations?Organizations gain peace of mind that large classes of attack on cloud guests are mitigated by Confidential Computing. Memory encryption with hardware key management and attestation prevents a compromise of the hypervisor becoming a compromise of guest data or integrity. Customers can now consider GCP as secure as private infrastructure for a much wider class of workloads. Canonical Ubuntu fully supports Confidential Computing on Google Cloud, providing a new level of trust in public cloud infrastructure.Burzin Patel, Vice President of Global Alliances, HashiCorpHashiCorp Vault enables teams to securely store and tightly control access to tokens, passwords, certificates, and encryption keys for protecting machines and applications. When combined with GCP’s Confidential Computing capabilities, confidentiality can be extended to the HashiCorp Vault server’s system memory, ensuring that malware, malicious privileged users, or zero days on the host cannot compromise data. Why did you choose Google Cloud as a partner for Confidential Computing?Google Cloud’s Confidential Computing nodes operate exactly like regular compute nodes making the offering very easy to use. We were able to take our existing Vault binary and host it on the Confidential Computing node to leverage the confidential computing benefits. No code or configuration changes were needed.What is the gap confidential computing solves specifically for your customers?Vault stores all of its sensitive data in memory and is stored as plaintext. In the past there were no easy solutions to keep this runtime memory protected. However, with the availability of confidential computing nodes, the data in memory is protected via encryption by utilizing the security features of modern CPUs together with confidential computing services.Any use-cases that are top of mind for you when it comes to confidential computing?HashiCorp Vault allows organizations to eliminate system complexity where any mistakes or misconfiguration could lead to a breach or data-leakage that in turn can halt operations and erode trust across customers. Together, HashiCorp Vault and Google Cloud’s Confidential Computing help organizations manage their most critical secrets and assets. This includes the entire secret lifecycle, from the initial creation, to sharing and distribution, and to the revocation or expiration of credentials and secrets.Any final thoughts?Security is the most critical element for enterprise customers looking to adopt the cloud. Customers are looking for a flexible solution that is robust and highly secure. The combination of HashiCorp Vault and Google Cloud Confidential Computing provide users a critical solution for their enterprise-wide cloud security needs.Mike Bursell, Chief Security Architect, Red HatAs more businesses and organizations move to the cloud, security remains a top priority. Maintaining the same levels of confidentiality that their partners, customers, regulators and shareholders expect across private and public clouds is vital. Red Hat believes that Confidential Computing is one key approach to extend security from on-premises deployments into the cloud, and Google’s announcement of Confidential VMs is an example of how customers can further secure their applications and workloads.What has Red Hat’s approach been to Confidential Computing?Red Hat Enterprise Linux is an enterprise operating system designed to handle the needs of customers across on-premises and hybrid cloud environments. Customers need stability, predictability and management solutions that scale with their workloads, which is why we enable Confidential Computing solutions in our product portfolio. That way customers don’t have to worry about migration costs. How will confidential computing impact cloud adoption?Often, customers with regulatory concerns have greater concerns about shifting into a truly open hybrid cloud environment, as they cannot expose their more sensitive data and applications outside their own data centers. Red Hat believes Confidential Computing can help them make this shift, expanding their opportunities for digital transformation, allowing them to provide quicker, more scalable and more competitive solutions, while maintaining the data privacy and protection assurances that their customers expect and require. As organizations balance the need for security with the opportunities presented by the cloud, Confidential Computing provides new ways to safely and securely embrace those opportunities.Dr. Thomas Di Giacomo, Chief Technology & Product Officer, SUSEConfidential VMs is a cloud industry security game-changer. This offering for our joint cloud customers expands sensitive data protection and compliance requirements, especially for regulated industries. The best part is you can run legacy and cloud-native workloads securely without any refactoring to the underlying application code, simplifying the transition to the cloud, all with little to no performance penalty.How has SUSE been working with Google Cloud and AMD?Working closely with AMD, SUSE added upstream support for AMD EPYC SEV processor to the Linux Kernel and was the first to announce Confidential VM support in SUSE Linux Enterprise Server 15 SP1 available in the Google Cloud Marketplace. These innovations allow our customers to take advantage of the scale and cost savings of Google Cloud Platform and the mission-critical manageability, compliance, and support from the #1 rated Linux support team, SUSE.How do you foresee this benefiting organizations?Confidential VMs will help tremendously accelerate our customer migrations to the cloud on their hybrid cloud digital transformation journey. This technology opens up new areas of migration opportunities for legacy on-premises workloads, custom applications as well as Private and Government workloads that require the utmost security and compliance requirements once considered not cloud-ready in the past.Solomon Cates, Principal Technologist, CTO Office, ThalesConfidential computing is a fundamental step in providing users control of their data as it goes “off premise” into cloud environments and all the way to the edge. Customers can essentially transition their workloads to the cloud with high assurance that includes auditable “proof” of control. And, architecturally, it opens up so many possibilities for customers.Many enterprises have significant trepidation when it comes to security in the cloud. Confidential computing helps alleviate that. For example, security professionals no longer have to worry about a cloud provider seeing or using their data.How does Confidential Computing help your customers?Confidential computing solves an issue that enterprises specifically have around trust in memory—namely that memory cannot be seen or used by a cloud provider. Three key use cases that can immediately benefit from this technology include edge computing, external key management and in-memory secrets.What made you partner with Google Cloud?Thales and Google Cloud have collaborated across a number of areas including cloud, security, Kubernetes containers and new technologies such as Continuous Access Evaluation Protocol (CAEP).  At the core, we both strive to offer customers the best option for strong security and privacy protection.Any final thoughts?From both a strategic and technical standpoint, Thales and Google Cloud have a shared vision that focuses on customer control and security of their data in the cloud. Through our work around confidential computing, we will bring new possibilities for securing workloads at the edge. Together, we are making it possible for enterprises to put their trust in the cloud with more sovereign control over their data security.We thank our hardware and software partners for their continuous innovation in this space. Confidential Computing can help organizations ensure the confidentiality of sensitive, business critical information and workloads, and we are excited to see the possibilities this technology will open up for your organization.
Quelle: Google Cloud Platform

Migrating apps to containers? Why Migrate for Anthos is your best bet

Most of us know that there is real value to be had in modernizing workloads, and there are plenty of customer success stories to showcase that. But, even though the value in modernizing workloads to Kubernetes has been well documented, there are still plenty of businesses that haven’t been able to make the jump. Reluctant businesses say that manually modernizing traditional workloads running on VMs into containers is a very complex/challenging project that involves significant time and costs. For instance, some proposals to refactor a single small to medium application can be $100,000 or more. Multiply that by 500 applications, and that’s a $50,000,000 project! To say nothing of how long it might take. Moreover, for some workloads (e.g., from third parties or ISVs) there is no access to the source code, precluding manual containerization altogether. As a result, these become blockers for many enterprises in their data center migration, especially for customers that don’t just want to lift and shift their important workloads. However, there’s an alternative. By leveraging automated containerization technologies and the right solution partners, you can cut the time and cost of a modernization project by as much as 90%, while enjoying most of the benefits that come with manual refactoring. Given that, using tools like Migrate for Anthos is a uniquely smart, efficient way to modernize traditional applications away from virtual machines and into native containers. Our unique automation approach extracts critical application elements from a VM so you can easily insert those elements into containers running on Google Kubernetes Engine (GKE), without artifacts like guest OS layers that VMs need but that are unnecessary for containers. For example, Migrate for Anthos automatically generates a container image, a Dockerfile for day-2 image updates and application revisions, Kubernetes deployment YAMLs and (where relevant) a persistent data volume onto which the application data files and persistent state are copied. This automated, intelligent extraction is significantly faster and easier than manually modernizing the app, especially when source code or deep application rebuild knowledge is unavailable. That’s why using Migrate for Anthos is one of the most scalable approaches to modernize applications with Kubernetes orchestration, image-based container management and DevOps automation.One of our customers, British newspaper The Telegraph, used Migrate for Anthos to accelerate its modernization and avoid the blockers we mentioned above. Here’s what Andrew Gregory, Systems Engineer Manager and Amit Lalani, Sr. Systems Engineer, had to say about the effort: “The Telegraph was running a legacy content management system (CMS) in another public cloud on several instances. Upgrading the actual system or migrating the content to our main Website CMS was problematic, but we wanted to migrate it from the public cloud it was on. With the help of our partners at Claranet and Google engineers, Migrate for Anthos delivered results quickly and efficiently. This legacy (but very important) system is now safely in GKE and joins its more modern counterparts, and is already seeing significant savings on infrastructure and reduced day-to-day operational costs.”Like at The Telegraph, any means that can accelerate and enable modernization of enterprise workloads is of high business value to our customers. Migrate for Anthos can accelerate and simplify the transition from VMs to GKE and Anthos by automating the containerization and “kubernetization” of the workloads. While manual refactoring typically takes many weeks or months, Migrate for Anthos can deliver containerization in hours or days. And once you’ve done so, you’ll start seeing immediate benefits in terms of infrastructure efficiency, operational productivity, and developer experience. To showcase that, Forrester’s New Technology Projection: The Total Economic Impact™ of Anthos (2019) report states: “When you are ready to migrate existing applications to the cloud, Migrate for Anthos makes that process simple and fast. The composite organization is projected to have 58% to 75% faster app migration and modernization process when using Anthos. After you containerize your existing applications you can take advantage of Anthos GKE, both on-prem and in the cloud, and consistently manage your Kubernetes deployments.”Let’s take a deeper look at some of the benefits you can expect from modernizing your VM-based workloads into containers on Kubernetes with Migrate for Anthos. Infrastructure efficiencyInternal Google studies have shown that converting VMs to containers in Kubernetes can yield between 30 – 65% savings on what you’re currently paying for your infrastructure, by means of: Higher utilization and density, leveraging automatic bin-packing and auto-scaling capabilities, Kubernetes places containers optimally in nodes based on required resources while scaling as needed, without impairing availability. In addition, unlike VMs, all containers on a single node share one copy of the operating system and don’t each require their own OS image and vCPU, resulting in a much smaller memory footprint and CPU needs. This means more workloads running on fewer compute resources. Shortened provisioning means you’re paying less to run the same workloads on account of them being ready sooner/easier. Operational productivityEmpowering your IT team to do more in less time also yields about 20 – 55% cost savings through reduced overall IT management and administration, for example:  Simplified OS management – In Anthos, the node and its operating system are managed by the system, so you don’t need to manage or be responsible for kernel security patches and upgrades. Configuration encapsulation – By leveraging declarative specification (infrastructure as code) you can simplify and automate your deployment and more easily perform maintenance tasks like rollback and upgrades. This all leads to a faster, more agile IT lifecycle.Reduced downtime – By leveraging Kubernetes features like self healing and dynamic scaling you’ll reduce incidents and have easier desired state management.Unified management – By modernizing legacy workloads into containers, DevOps engineers can use the same method to manage all their workloads, both cloud-native and cloud “naturalized” workloads, making it faster and easier for IT to manage your hybrid IT landscape. Environment parity with improved visibility and monitoring, makes finding and fixing problems less toilsome. Developer productivity When you’ve got a better and more agile IT environment, your developers can do more with less, usually resulting in cost savings from developer efficiency and reduced infrastructure. Apps that have been converted into containers benefit from:Layering efficiency – The ability to use Docker images and layers (which Migrate for Anthos extracts as part of the container artifacts). Developer velocity – You can finally “write once run everywhere,” and combine automated CI/CD pipelines with on-demand, repeatable test deployments using declarative models and Kubernetes orchestration.Faster lifecycle – Get products to market quicker, yielding additional revenue and competitive market advantages, on top of savings. In short, modernizing your VMs into containers running on Kubernetes has benefits across infrastructure, operations, and development. Although modernization may seem intimidating at first, Migrate for Anthos helps make this process fast and painless. You can read more about it here, watch a quick video on using Migrate for Anthos on Linux or Windows workloads, or try it yourself using Qwiklabs.And if you’re interested in talking to someone about using Migrate for Anthos please fill out this form (mention “Migrate for Anthos” in the ‘Your Project’ field) and someone will contact you directly.
Quelle: Google Cloud Platform

Dataproc Metastore: Fully managed Hive metastore now in public preview

The Apache Hive metastore service has become a building block for data lakes that utilize the diverse world of open-source software, such as Apache Spark and Presto. We’re launching the Dataproc Metastore into public preview today, so these powerful tools are now easy to use by any Google Cloud customer with fewer distractions and delays. The Dataproc Metastore is a fully managed, highly available, auto-healing, open-source Apache Hive metastore service that simplifies technical metadata management for customers building data lakes on Google Cloud. And for a limited time only, it’s free! This launch exemplifies our commitment to fast-paced innovation and delivery, combining cloud technology with open source, and closely follows our announcement of the private preview in June of this year. Before we go into more detail, we would also like to thank our private preview users for testing and providing rich feedback—the launch today has been made better with your valuable input.  What does this mean for my data lake?If you are familiar with the Hive Metastore, you likely already know it is a critical component of many data lakes because it acts as a central repository of metadata. In fact, a whole ecosystem of tools, open-source and otherwise, are built around the Hive Metastore, some of which this diagram illustrates.The Dataproc Metastore is a serverless Hive Metastore that unlocks several key data lake use cases in Google Cloud, including:Many ephemeral Dataproc clusters can utilize a Dataproc Metastore at the same time, allowing many users of open-source tools, such as Spark , Hive, and Presto, to access consistent metadata at the same time. Unifying metadata between open-source tables and Data Fusion, so ETL and ELT on those tables is easier and code-free.Tying together metadata into a central store so cloud-natiove services like Dataproc can seamlessly interoperate with other open-source tools or partner technologies.The Dataproc Metastore now means your data lake is easier to manage, more unified, and increasingly serverless for fewer distractions. New features in Dataproc MetastoreThroughout the private preview period, and since our initial announcement in June, we have added many new features to the Dataproc Metastore. Several of these new features are launching with this release today.IAM and Kerberos—Fine-grained Cloud Identity and Access Management (Cloud IAM) support, along with out-of-the-box support for Kerberos and other security tools such as Apache Ranger.Import/export—Metadata can be imported and exported to enable bidirectional integration with and migration from other Hive Metastores, such as those on-premises.VPC-SC—Support for Google Cloud VPC Service Controls to mitigate data exfiltration risks.ACID transactions—Dataproc Metastore supports ACID transactions using Hive’s ACID transaction capabilities.Cloud Monitoring integration—Logging and monitoring of Dataproc Metastore instances seamlessly inside of Cloud Monitoring and Logging.Broad Dataproc compatibility—Compatible with a broad range of Dataproc releases, including the Dataproc 2.0 preview release with Spark, Hadoop, and Hive 3.x. Service updates—You can transactionally update elements of the hive Metastore service including configurations, tiers, ports, maintenance window, and more.Cloud Console and Cloud SDK—Dataproc Metastore supports both the Cloud Console and the Cloud SDK command line (gcloud beta metastore).We will continue to move quickly to get the Dataproc Metastore into general availability while also adding highly requested features such as customer-managed encryption keys.Dataproc Metastore public preview pricingDuring the public preview period, which starts today and lasts until GA, the Dataproc Metastore will be offered at a 100% discount. This discount is intended for you to use and test the technology without incurring costs for the testing. The Dataproc Metastore is offered in two service tiers, developer and enterprise, each of which offer different features, service levels, and pricing because they are intended for different use cases.This pricing allows you to create developer instances for quick testing and prototyping without needing to test against your production environment or create multiple copies of your production database. The enterprise tier is intended for production deployments that require high availability, performance, and stability. Future releases will also incorporate features targeted at specific tiers, such as Data Catalog integration.You can find more information in the pricing documentation for Dataproc Metastore.Serverless open sourceThe Dataproc Metastore is a good example of how the best of Google Cloud infrastructure can be used to run managed open source. As a result of innovations in how we run, secure, and scale the Hive Metastore, we have been able to make the Dataproc Metastore serverless. This launch is the beginning of how we’re reshaping managed open source for data analytics in cloud. As a team passionate about both cloud and open source, it is our goal to bring the very things that make the Hive Metastore uniquely great, including no infrastructure to manage, automated scalability, enhanced hands-off high availability, and easier pricing to other popular open source components in the future. Get startedAny Google Cloud customer can use the Dataproc Metastore, for free during preview, starting today. You can follow the quickstart guide or review the full documentation for more information on how to get started.Related ArticleDataproc Hub makes notebooks easier to use for machine learningDataproc Hub, now generally available, makes it easy to use open source, notebook-based machine learning on Google Cloud, powered by Spark.Read Article
Quelle: Google Cloud Platform

Create a secure and code-free data pipeline in minutes using Cloud Data Fusion

Organizations are increasingly investing in modern cloud warehouses and data lake solutions to augment analytics environments and improve business decisions. The business value of such repositories increases as additional data is added. And with today’s connected world and many companies adopting a multi-cloud strategy, it is very common to see a scenario where the source data is stored in a cloud provider different from where the final data lake or warehouse is deployed. Source data may be in Azure or Amazon Web Services (AWS) storage, for example, while the data warehouse solution is deployed in Google Cloud. Additionally, in many cases, regulatory compliance may dictate the need to anonymize pieces of the content prior to loading it into the lake so that the data can be de-identified prior to data scientists or analytic tools’ consumption. Last, it may be important for customers to perform a straight join on data coming from disparate data sources and apply machine learning predictions to the overall dataset once the data lands in the data warehouse. In this post, we’ll describe how you can set up a secure and no-code data pipeline and demonstrate how Google Cloud can help you move data easily, while anonymizing it in your target warehouse. This intuitive drag-and-drop solution is based on pre-built connectors, and the self-service model of code-free data integration removes technical expertise-based bottlenecks and accelerates time to insight. Additionally, this serverless approach that uses the scalability and reliability of Google services means you get the best of data integration capabilities with a lower total cost of ownership.Here’s what that architecture will look like:Understanding a common data pipeline use caseTo provide a little bit more context, here is an illustrative (and common) use case:An application is hosted at AWS and generates log files on a recursive basis. The files are compressed using gzip and stored on an S3 bucket. An organization is building a modern data lake and/or cloud data warehouse solution using Google Cloud services and must ingest the log data stored in AWS.The ingested data needs to be analyzed by SQL-based analytics tools and also be available as raw files for backup and retention purposes. The source files contain PII data, so parts of the content need to be masked prior to its consumption.New log data needs to be loaded at the end of each day so next day analysis can be performed on it. Customer needs to perform a straight join on data coming from disparate data sources and apply machine learning predictions to the overall dataset once the data lands in the data warehouse. Google Cloud to the rescueTo address the ETL (extract,transform and load) scenario above, we will be demonstrating the usage of four Google Cloud services: Cloud Data Fusion, Cloud Data Loss Prevention (DLP), Google Cloud Storage, and BigQuery. Data Fusion is a fully managed, cloud-native, enterprise data integration service for quickly building and managing data pipelines. Data Fusion’s web UI allows organizations to build scalable data integration solutions to clean, prepare, blend, transfer, and transform data without having to manage the underlying infrastructure. Its integration with Google Cloud simplifies data security and ensures data is immediately available for analysis. For this exercise, Data Fusion will be used to orchestrate the entire data ingestion pipeline. Cloud DLP can be natively called via APIs within Data Fusion pipelines. As a fully managed service, Cloud DLP is designed to help organizations discover, classify, and protect their most sensitive data. With over 120 built-in InfoTypes, Cloud DLP has native support for scanning and classifying sensitive data in Cloud Storage and BigQuery, and a streaming content API to enable support for additional data sources, custom workloads, and applications. For this exercise, Cloud DLP will be used to mask sensitive personally identifiable information (PII) such as a phone number listed in the records. Once data is de-identified, it will need to be stored and available for analysis in Google Cloud. To cover the specific requirements listed earlier, we will demonstrate the usage of Cloud Storage (Google’s highly durable and geo-redundant object storage) and BigQuery, Google’s serverless, highly scalable, and cost-effective multi-cloud data warehouse solution. Conceptual data pipeline overviewHere’s a look at the data pipeline we’ll be creating that starts at an AWS S3 instance, uses Wrangler and Redact API for anonymization, and then moves data into both Cloud Storage or BigQuery.Walking through the data pipeline development/deployment processTo illustrate the entire data pipeline development and deployment process, we’ve created a set of seven videos. You’ll see the related video in each of the steps here. Step 1 (Optional): Did not understand the use case yet or would like to watch a refresh? This video provides an overview of the use case, covering the specific requirements to be addressed. Feel free to watch it if required.Step 2: This next video covers how the source data is organized. After watching the recording, you will be able to understand how the data is stored in AWS and explore the structure of the sample file used by the ingestion pipeline.Step 3: Now that you understand the use case goals and how the source data is structured, start the pipeline creation by watching this video. On this recording you will get a quick overview of Cloud Data Fusion, understand how to perform no-code data transformations using the Data Fusion Wrangler feature, and initiate the ingestion pipeline creation from within the Wrangler screen.Step 4: As mentioned previously, de-identifying the data prior to its consumption is a key requirement of this example use case. Continue the pipeline creation and understand how to initiate Cloud DLP API calls from within Data Fusion, allowing you to perform data redaction on the fly prior to storing it permanently. Watch this video for the detailed steps.Step 5: Since the data is now de-identified, it’s time to store it in Google Cloud. Since the use case mandated both structured file backups and SQL-based analytics, we will store the data in both Cloud Storage and BigQuery. Learn how to add both Cloud Storage and BigQuery sinks to the existing pipeline in this recording.Step 6: You are really close now! It’s time to validate your great work. Wouldn’t it be nice to “try” your pipeline prior to fully deploying it? That’s what the pipeline preview feature allows you to do. Watch this quick video and understand how to preview and subsequently deploy your data ingestion pipeline, taking some time to observe the scheduling and deployment profile options.Step 7: Woohoo! Last step. Check this video out and observe the ability to analyze the full pipeline execution. In addition, this recording will cover how to perform high-level data validation on both Cloud Storage and BigQuery targets.Next steps:Have a similar challenge? Try Google Cloud and this Cloud Data Fusion quickstart next. Have fun exploring!
Quelle: Google Cloud Platform

The serverless gambit: Building ChessMsgs.com on Cloud Run

While watching The Queen’s Gambit on Netflix just recently, I was reminded about how much I used to enjoy playing chess. I was eager to play a game, so I started to tweet, “D2-D4” knowing that someone would recognize this as an opening move and likely respond with their move, giving me the fix I needed. I paused before hitting the tweet button because I realized that I’d need to set up a board (physical or virtual) to keep track of the game. If I received multiple responses, I’d need multiple boards. I decided not to send the tweet.Later in the day, I had the idea to create a simple service that addresses my use case. Instead of designing a full chess site, I decided to create a chess board logger/visualizer to make it practical to play via Twitter or any other messaging/social platform.Instead of tweeting moves back and forth, players tweet links back and forth, and those links go to a site that renders the current chessboard, allows a new move, and creates a new link to paste back to the opponent. I wanted this to be 100% serverless, meaning that it will scale to zero and have zero maintenance requirements. Excited about this idea, I put together a shopping list:My MVP requirements:Represent the board position—ideally completely in the URL to keep it stateless from a server-side perspectiveDisplay a chessboard and let the player make their next move.Stretch goals:Enforce chess rules (allow only legal moves).Dynamically create a png/jpg of the chessboard that I can use as an Open Graph and Twitter card image so that when a player sends the link, the image of the board will automatically display.Putting it all togetherRepresenting the board positionThere is a standard notation for describing a particular board position of a chess game called Forsyth–Edwards Notation (FEN) that was exactly what I needed. A FEN is a sequence of ASCII characters. For example, the starting position for any chess game can be represented by the following string:rnbqkbnr/pppppppp/8/8/8/8/PPPPPPPP/RNBQKBNR w KQkq – 0 1Each letter is a piece: pawn = “P”, knight = “N”, bishop = “B”, rook = “R”, queen = “Q” and king = “K”. Uppercase letters represent white pieces and lowercase letters represent black. The last part of the string is specific to certain rules in chess (read more about FEN).I knew I could use this in the URL, so my first requirement was complete and I was able to represent the board state in the URL eliminating the need for a backend data store.Displaying the chessboard and allowing drag-and-drop movesNumerous chess libraries are available. One in particular that caught my eye was chessboard.js—described as “a JavaScript chessboard component with a flexible ‘just a board’ API”. I quickly discovered that this library can display chess boards from a FEN, allow pieces to be moved, and update the FEN. Perfect!In only two hours, I had the basic functionality implemented.Enforcing chess rulesI originally thought that making this service aware of chess rules would be difficult, but then I saw the example in the chessboard.js docs showing how to integrate it with another library called chess.js—“a JavaScript chess library that is used for chess move generation/validation, piece placement/movement, and check/checkmate/stalemate detection—basically everything but the AI”. A short time later, I had it working! Stretch goal #1 completed.Where’s what a couple of game moves look like:Moving the pawn from D2 to D4 in a new game—https://chessmsgs.com/?fen=rnbqkbnr%2Fpppppppp%2F8%2F8%2F3P4%2F8%2FPPP1PPPP%2FRNBQKBNR+b+KQkq+d3+0+1&to=d4&from=d2&gid=mOhlhRlMboYsHLqBF1f7IBlack countering with a similar move of pawn from D7 to D5—https://chessmsgs.com/?fen=rnbqkbnr%2Fppp1pppp%2F8%2F3p4%2F3P4%2F8%2FPPP1PPPP%2FRNBQKBNR+w+KQkq+d6+0+2&to=d5&from=d7&gid=mOhlhRlMboYsHLqBF1f7IThe URL has the following data:fen—the new board positionfrom and to—indicating what move occurred (I use this to highlight the squares)gid—a unique game ID (I used nanoid)—I’ll use this to connect moves to a single game in the future. For example, I could add a feature that lets the user request the entire game transcript). Done! Except…At this point, there were no server requirements other than simple HTML static hosting. But after playing it with some friends and family, I decided that I really wanted to accomplish the other stretch goal—dynamically create a png/jpg of the chessboard that I can use as an Open Graph and Twitter card image.  With this capability, an image of the board will automatically display when a player sends the link. Without it, the game is a series of ugly URLs.Dynamically creating the Open Graph imageThis requirement introduced some server-side requirements. I needed two things to happen on the server.First, I needed to dynamically generate a board image from a FEN. Once again, open source to the rescue (almost). I found chess-image-generator, a JavaScript library that creates a png from a FEN. I wrapped this in a bit of Node.js/Express code so that I could access the image as if it were static. For example, here’s a demo of the real endpoint: https://chessmsgs.com/fenimg/v1/rnbqkb1r/ppp1pppp/5n2/3p4/3P4/2N5/PPP1PPPP/R1BQKBNR w KQkq – 2 3.png. This link results in this image:Second, I needed to dynamically inject this FEN-embedded URL into the content attribute of the meta tag in the main HTML. Like me, you might be thinking that you could just do some DOM manipulation in JavaScript and avoid having to dynamically change HTML on the server. But, the Open Graph image is retrieved by a bot from whatever service you use for messaging. These bots don’t execute any client-side JavaScript and expect all values to be static. So, that led to additional server-side work.I needed to dynamically convert this:Into something like this:I could have used one of many Node templating engines to do this, but they all seemed like overkill for this simple substitution requirement, so I just wrote a few lines of code for some string.replace() calls in my Node server. With this functionality added, a game on Twitter (and other services) now looks much better:Check out the codeThe source for chessmsgs.com is available on GitHub at https://github.com/gregsramblings/chessmsgs. Deciding where to host itThe hosting requirements are simple. I needed support for Node.js/Express, domain mapping, and SSL. There are several options on Google Cloud including Compute Engine (VMs), App Engine, and Kubernetes Engine. For this app, however, I wanted to go completely serverless, which quickly led me to Cloud Run. Cloud Run is a managed platform that enables you to run stateless containers that are invocable via web requests or Pub/Sub events. Cloud Run is also basically free for this type of project because the always-free-tier includes 180,000 vCPU-seconds, 360,000 GiB-seconds, and 2 million requests per month (as of this writing—see the Cloud Run pricing page for the latest details). Even beyond the free tier, it’s very inexpensive for this type of app because you only pay while a request is being handled on your container instance, and my code is simple and fast.Lastly, deploying this on Cloud Run brings a lot of added benefits such as continuous deployment via Cloud Build, and log management and analysis via Cloud Logging, both of which are super easy to set up.What’s next?If this suddenly becomes the most popular site of the day, I’m actually in good shape from a scalability point of view because of my decision to use Cloud Run. If I really wanted to engineer this for extreme loads, I could easily deploy it to multiple regions throughout the globe and set up a load balancer and possibly a CDN. I also could separate the web hosting functionality from the image generation functionality to allow each to scale as needed.When I first started thinking about the image generation, I naturally thought about caching the images in Google Cloud Storage. This would be easy to do and storage is crazy cheap. But, then I did a bit of research and learned the following fun facts. After two moves (one move for each player), there are 400 different distinct board positions. After each player moves again (two moves each), this number is now 71,782 distinct positions. After each player moves again (three moves each), the number is now 9,132,484 distinct positions! I could gain a bit of performance by caching the most popular openings, but each game would quickly go beyond the cached images so it didn’t seem worth it. By the way, to cache every possible board position would be about 1046 positions, which is a massive number that doesn’t even have a name.ConclusionThis was a fun project – almost therapeutic for me since my “day job” doesn’t allow much time for writing code. If this becomes popular, I’m sure others will have ideas on how to improve it. This was my first hands-on with Cloud Run beyond the excellent Quick Starts (examples for Go, Node.js, Python, Java, C#, C++, PHP, Ruby, Shell, etc.). Because of my role in developer advocacy at Google, I was aware of most Cloud Run capabilities and features but after using it for something real, I now understand why developers love it!Where to learn moreCloud Run Product PageCloud Run DocsHello Cloud Run QwiklabThe Cloud Run unofficial FAQ (created by co-worker Ahmet Alp Balkan and community maintained)The Cloud Run Button—Add a click-to-deploy button to your git reposList of Cloud Run videos from our YouTube channelNEW O’Reilly Book: Building Serverless Applications with Google Cloud Run by Wietse VenemaAwesome Cloud Run—massive curated list of resources by Steren Giannini‎ (Cloud Run PM)Related Article3 cool Cloud Run features that developers loveCloud Run developers enjoy pay-per-use pricing, multiple concurrency and secure event processing.Read Article
Quelle: Google Cloud Platform

Download and Try the Tech Preview of Docker Desktop for M1

Last week, during the Docker Community All Hands, we announced the availability of a developer preview build of Docker Desktop for Macs running on M1 through the Docker Developer Preview Program. We already have more than 1,000 people testing these builds as of today. If you’re interested in joining the program for future releases you should do it today!

As I’m sure you know by now, Apple has recently shipped the first Macs based on the new Apple M1 chips. Last month my colleague Ben shared our roadmap for building a Docker Desktop that runs on this new hardware. And I’m delighted to tell you that today we have a public preview that you can download and try out.

Like many of you, we at Docker have been super excited to receive and code with these new computers: they just feel so fast! We also know that Docker Desktop is a key part of the development cycle for over 3M developers using Docker Desktop with over half of you on Macs. To support all our Mac users we’ve been working hard to get Docker Desktop ready to run on the new M1 hardware. It is not release quality yet, or even beta quality, but we have an early preview build and we wanted to let you try it as soon as possible.

How We Got to a Technical Preview

When Ben announced that we were working on adapting Docker Desktop on this new hardware. We had roughly 3 engineering challenges to tackle to get this release out to you: 

Migrate from HyperKit to the Virtualization Framework.

One of the key challenges for the Docker Desktop team was to replace HyperKit, which Docker open sourced back in 2016, with the Virtualization Framework provided by Apple which was included in macOS Big Sur.

Recompile all the various binaries of Docker Desktop in native arm.

Many of the tools that we use in our toolchain to build these binaries are not yet ready to support the M1 Mac as of today. At Docker, we use the Go language extensively, and Docker Desktop is no exception. The Go language will support Apple Silicon in their 1.16 release which is targeted for February 2021.

Have enough hardware to reliably run continuous deployment on M1 macs.

The Docker Desktop team relies heavily on automated testing through continuous integration to ensure the quality of our releases. Until this week our continuous integration could not be set up because none of our partners had enough M1 machines yet. Fortunately, we are working with MacStadium and we are setting up new M1 Macs on our CI system.

Thanks to the significant progress we have been able to make on the first two steps, we are sharing a Tech Preview of Docker Desktop for M1 today. Download it here!

Multi-Platform Baked In

Many developers are going to experience multi-platform development for the first time with the M1 Macs. This is one of the key areas where Docker shines. Docker has had support for multi-platform images for a long time, meaning that you can build and run both x86 and ARM images on Desktop today. The new Docker Desktop on M1 is no exception; you can build and run images for both x86 and Arm architectures without having to set up a complex cross-compilation development environment.

Docker Hub also makes it easy to identify and share repositories that provide multi-platform images.

And finally, using docker buildx you can also easily integrate multi-platform builds into your build pipeline.

Try the M1 Preview Today

Right on time for the year-end festivities, we’re excited to share with you our M1 Preview:

Here is the Download!

Keep in mind that this is a preview release: it may break, it has not been tested as thoroughly as our normal releases and ‘here be dragons’. Your help is needed to test Docker Desktop on Apple Silicon so that we can continue to provide a great developer experience on all Apple devices. You can help us by providing bug reports on docker/for-mac. We will use this feedback to help us improve and iterate on both the Desktop product and the multi-architecture experience as we aim to provide a GA build of Docker Desktop in the first quarter of 2021.

In the meantime, enjoy this tech preview build of Docker Desktop for M1. Happy Holidays!
The post Download and Try the Tech Preview of Docker Desktop for M1 appeared first on Docker Blog.
Quelle: https://blog.docker.com/feed/