Unlock the power of change data capture and replication with new, serverless Datastream, now GA

We’re excited to announce that Datastream, Google Cloud’s serverless change data capture (CDC) and replication service, is now generally available. Datastream allows you to synchronize data across disparate databases, storage systems, and applications reliably and with minimal latency to support real-time analytics, database replication, and event-driven architectures. You can easily and seamlessly deliver change streams from Oracle and MySQL databases into Google Cloud services such as BigQuery, Cloud SQL, Google Cloud Storage and Cloud Spanner, saving time and resources and ensuring your data is accurate and up to date. Get started with Datastream today.Datastream provides an integrated solution for CDC replication use cases with custom sources and destinations*Check the documentation page for all supported sources and destinations.Since our public preview launch earlier this year, we’ve seen Datastream used across a variety of industries, by customers such as Chess.com, Cogeco, Schnuck Markets, and MuchBetter. This early adoption strengthens the message we’ve been hearing from customers about the demand for change data capture to provide replication and streaming capabilities for real-time analytics and business operations. MuchBetter is a multi-award-winning e-wallet app, providing a truly secure and enjoyable banking alternative for customers all over the world. Working with Google Cloud Premier Partner Datatonic, they’re leveraging Datastream to replicate real-time data from MySQL OLTP databases into a BigQuery data warehouse to power their analytics needs. According to Andrew McBrearty, Head of Technology at MuchBetter, “from MuchBetter’s point of view, leveraging Dataflow, BigQuery and Looker has unlocked additional insights from our ever-increasing operational data. Using Datastream in our solution ensured continued real-time capability – we now have trend analysis in place, improved efficiency across the business, and the ability to use our data to derive actionable insights and to make data-driven decisions. This means we can continue to grow and adapt at a pace our customers have come to expect from MuchBetter. And for the first time, the world of ML and AI is open to us.”Getting to know DatastreamGoogle Cloud customers are choosing Datastream for real-time change data capture because of its differentiated approach:Simple experienceReal-time replication of change data shouldn’t be complicated: database preparation documentation, secure connectivity setup, and stream validation should be built right into the flow. Datastream delivers on this experience, as MuchBetter discovered during their evaluation of the product. “Datastream’s ease-of-use and immediate availability (serverless) meant we could start our evaluation and immediately see results”, says Mark Venables, Principal Data Engineer at MuchBetter. “For us, this meant getting rid of the considerable pre-work needed to align proof of concept tests with third-party CDC suppliers.”Datastream guides you to success by providing detailed pre-requisites and step-by-step configuration guidelines to prepare your source database for CDC ingestion.End-to-end solutionBuilding pipelines to replicate changes from your source database shouldn’t take up all of your team’s time. Use pre-built Dataflow templates to easily replicate data into BigQuery, Cloud Spanner or Cloud SQL. Out of the box, these Dataflow templates will automatically create the tables and update the data at the destination, taking care of any out-of-order or duplicate events, and providing error resolution capabilities. Leverage the templates’ flexibility to fine-tune Dataflow to fit your specific needs. “Google-managed Dataflow templates meant getting our pipelines up and running with minimal effort and fuss – this allowed more time to be spent on more complex pipeline development whilst tactically delivering solutions to our users,” says Venables.Secure Datastream keeps your migrated data secure, supporting private connectivity between source and destination databases. “Establishing connectivity is often viewed as hard. Datastream surprised us with its ease of use & setup, even in more secure modes,” says Grzegorz Dlugolecki, Principal Cloud Architect at Chess.com, a leading online chess community and mobile application, hosting more than ten million chess games every day. “Datastream’s private connectivity configuration allowed us to easily create a private connection between our source and the destination, and ensure our data is safe and secure.”Datastream provides a simple wizard to automatically set up private, secure connectivity to your source databaseHigh throughput, low latencyWith Datastream’s serverless architecture, you don’t need to worry about provisioning, managing machines, or scaling up resources to meet fluctuations in data throughput. Datastream guarantees high performance – a single stream can process 10’s of MBs per second, while ensuring minimal latency. “We evaluated several market-leading ETL solutions”, says Dlugolecki,  “Datastream was the only tool able to successfully sync our complex, single-table datasets, doing this in weeks instead of years estimated by the other vendors.”Getting started with DatastreamYou can start streaming real-time changes from your Oracle and MySQL databases today using Datastream:Navigate to the Datastream area of your Google Cloud console, under Big Data, and click Create Stream.Choose the source database type, and see what actions you need to take to set up your source.Create your source connection profile, which can later be used for additional streams.Define how you want to connect your source.Create and configure your destination connection profile.Validate your stream and make sure the test was successful. Start the stream when you’re ready.Once the stream is started, Datastream will backfill historical data and will continuously replicate new changes as they happen. Learn more and start using Datastream todayDatastream is now generally available for Oracle and MySQL sources. Datastream supports sources both on-premises and in the cloud, and captures historical data and changes into Cloud Storage. Integrations with Cloud Data Fusion and Cloud Dataflow (our data integration and stream processing products, respectively) replicate changes to other Google Cloud destinations, including: BigQuery, Cloud Spanner, and Cloud SQL.For more information, head on over to the Datastream documentation, see our step-by-step Datastream + Dataflow to BigQuery tutorial, or start training with this Datastream Qwiklab.Related ArticleUsing Datastream to unify data for machine learning and analyticsWhile machine learning model architectures are becoming more sophisticated and effective, the availability of high-quality, fresh data fo…Read Article
Quelle: Google Cloud Platform

How data and AI can help media companies better personalize; and what to watch out for

Media companies now have access to an ever-expanding pool of data from the digitally connected consumer. And over the past two years, as content consumption and audience behaviors have shifted in response to the world around us, direct-to-consumer has only accelerated. As media organizations pivot from third-party to first-party data, this presents challenges with the volume, velocity and fragmentation of data. It’s also an opportunity to better understand how to acquire, engage and retain audiences ⁠— and inject agility into their business amidst a competitive landscape. How should media companies be thinking about their data, and its value, to capitalize on this opportunity?  To help answer these questions, we sat down with Gloria Lee, Executive Account Director in Media & Entertainment and John Abel, Technical Director for the Office of the CTO at Google Cloud.Data and the growing importance of personalization There’s no doubt customer needs and expectations are in a constant state of flux. Across the media industry, audiences are increasingly expecting personalized content. In fact, a  PWC study conducted in 2020 found that nearly one-third (31%) of survey respondents said easy, personalized content recommendations would be a reason for staying with a streaming service.Audience engagement is the currency, and in a crowded space where attention is finite, media companies need a granular understanding of their audience. To do this, there’s an opportunity to capture and capitalize on first party data, so they can better serve their audience. Not just what audiences are consuming, but also when, where, and on which platform (and increasingly, those platforms are digital). These data points are key in understanding audiences deeply to deliver hyper-personalized experiences that audiences are expecting.    “If you look across the world today, we know that through digitalization, [that] hyper-personalization is required,” says John. “So that hyper-personalization, the volume of data and the value of the data is super critical across all industries. Media and entertainment is no different,” he adds.Enriching storytelling through AI & ML   Extracting insights on how, where and when the consumer wants to receive content will accelerate the need for data research; AI and ML will be critical to unlocking data’s full potential. “The most valuable data is generated data, typically from machine learning or AI, where you’re seeing new insights in data that give you new opportunities.” explains Gloria.  New technologies are providing insights — often in real time — about audiences, making personalization an easier task. An example use case would be recommending a new song based on a user’s listening history. This kind of personalization is just the start, as AI/ML unlocks more novel opportunities. For example, AL/ML can also be used to enrich the watching experience by finding opportune moments to integrate brands. As Gloria puts it, “Artificial intelligence, and machine learning is what enables people to quickly look through their content to find relevant moments for marketing purposes”. Getting personalization right, while making sure to keep consumer information safe and private is a challenge for all consumer companies; not just M&E. John explains “there’s a blend of how they move technology to the edge and they don’t break privacy.” Media and Entertainment companies will need to keep their data secure and private, using sophisticated practices like data federation. In this model, individual data is not exchanged.  Rather, data is first aggregated into cohorts to anonymize the individual. The goal of methods like this is to obtain useful insights while retaining privacy and security.How data is driving audience experiences Spotify is a prime example of a media company using data-led insights to provide personalized content for their customers — making it easier for users to discover new audio content and connect with their favorite artists or podcasts. “[With] Google Cloud…we can iterate quicker on key needs, like data insights and machine learning…[streamlining] our ability to concentrate on what’s important to our users and give them the experiences they know and love about Spotify.”—Tyson Singer, VP of technology and platform at SpotifySky, one of Europe’s  leading broadcasters is also transforming its data strategy to better serve their customers. By creating a scalable cloud-based architecture, Sky can keep up with increasing amounts of TV box diagnostic data on service uptime and delivery ⁠— meaning less data lost and more time to focus on improving user experience through personalization. “The data will sit right at the heart of Sky’s future strategy. It will help ensure that our products are intuitive and easy to use and that we can keep seamlessly connecting customers with the content and services they know and love,” says Oliver Tweedie, Director of Data Engineering at Sky.Transforming to a data-oriented Media companyKeeping up with new technology trends, inside and outside of the industry, will play a critical role in how media and entertainment companies can survive and thrive into the future. And without a way to centralize and draw insights from their data quickly, media organizations will struggle to stay in the race. With any type of change comes resistance. But at the end of the day, it all comes down to people. When navigating digital transformations, Gloria touches on the three categories of people: ⁠supporters, those excited about the change, those who couldn’t care less and detractors, those who are opposed to it. “It’s really tapping into the leaders for those three different groups within the company and trying to get them on board and seeing what their drivers are,” Gloria explains.So what advice do John and Gloria have for Media players looking into the data-led future?Related ArticleRead Article
Quelle: Google Cloud Platform

Expanding our infrastructure with cloud regions around the world

Businesses and organizations around the world depend on Google Cloud to help them digitally transform, innovate across their industries, and drive operational efficiencies for long-term growth. Our global network of Google Cloud Platform regions is the foundation of this capability, delivering high-performance, low latency cloud-based services to customers and their users in more than 200 countries and territories globally. With 29 cloud regions and 88 zones, we operate more regions with multiple availability zones than any other hyperscale cloud provider. So far in 2021, we’ve opened new regions in Warsaw (Poland), Delhi NCR (India), Melbourne (Australia), and Toronto (Canada), bringing the cleanest cloud in the industry closer to more customers across multiple continents. Building on this momentum, today we are excited to share further updates to our expansion strategy.Extending our Google Cloud region roadmapChileToday, we’re proud to announce that our Santiago cloud region is now operational, ready to help more South American customers and partners build a digital-first future. This marks our first cloud region in Chile and second in South America, complementing São Paulo, which opened in 2017. The Santiago cloud region brings our high-performance, low-latency cloud services closer to customers across Latin America, from financial institutions like Caja Los Andes to health providers like Red Salud and enterprises like LATAM Airlines. Join us as we celebrate the opening of the Santiago cloud region with our customers and local Google leaders. IsraelToday, we are excited to share that our Google Cloud region in Israel will be located near Tel Aviv. When operational, the Tel Aviv region will enable us to meet growing demand for cloud services in Israel across industries, from retail to financial services to the public sector.Israeli companies like BreezoMeter, Haaretz, PayBox, and Wix already run on our cloud, making it easier to operate their businesses faster, securely, and more reliably. “Google Cloud’s global network helps Wix to achieve the best performance around the globe. For example, with Google Cloud CDN, we are able to serve tens of millions of requests per day seamlessly, while ensuring that our customers get a consistently great web experience worldwide.” – Eugene Olshenbaum, VP Technology at WixRecently, Google Cloud was selected by the Israeli government to provide public cloud services to all government entities from across the state, including ministries, authorities, and government-owned companies.GermanyOur second cloud region in Germany will be located in Berlin-Brandenburg, complementing our existing cloud region in Frankfurt. Once launched, our cloud region in Berlin-Brandenburg will strengthen our safe and secure platform for customers in Germany, including both public sector organizations and businesses like BMG, helping them scale and adapt to changing requirements.”At BMG, we’re continuously pushing digitization further to help ensure that when our artists release new music, it is promoted effectively around the world. With autoscaling via BigQuery, excellent customer support, and a clean and simple user interface, Google Cloud has been a partner to our technology team and beyond. The new Google Cloud region in Berlin-Brandenburg will further improve collaboration company-wide and make data more accessible to all teams.” – Gaurav Mittal, VP IT & Systems at BMGSaudi ArabiaLast year, we announced our plans to deploy and operate a cloud region in Saudi Arabia, while a local strategic reseller, sponsored by Aramco, will offer cloud services to businesses in the Kingdom. Today, we are announcing Dammam as the location for this cloud region. As we prepare for launch, we will start hiring out of a Riyadh-based office to support the cloud region’s deployment and operation.  United StatesOver the next year, we will add cloud regions in Columbus, Ohio, and Dallas, Texas, providing customers operating in North America with the capacity they need to run mission-critical services at the lowest possible latency. These new U.S. regions will bring our services even closer to existing customers such as J.B. Hunt Transport, Inc., which is implementing Google Cloud solutions to help create the most efficient transportation network in North America.Beyond performance and capacityAs we add regions across the Americas, Asia, Europe and the Middle East, Google Cloud is committed to continuing to help build a more sustainable future and create opportunities for everyone. We operate the cleanest cloud in the industry. Google was the first organization of its size to become carbon neutral in 2007, and we were the first major company to match 100% of our electricity consumption with renewable energy starting in 2017 and every year since then. As we now work toward achieving carbon-free energy 24/7 by 2030, we’re proud to support our customers with cloud infrastructure and tools to reduce their environmental impact. In addition to decarbonizing our energy consumption around the world, we are committed to upholding human rights in every country where we operate. This includes respecting the Universal Declaration of Human Rights, as well as the standards established in the United Nations Guiding Principles on Business and Human Rights and Global Network Initiative Principles. We’re a proud founding member of the Global Network Initiative, in which we work closely with civil society, academics, investors and industry peers to protect and advance freedom of expression and privacy globally as we deliver high-quality, relevant and useful content.Whenever we expand operations in a new country, we undertake thorough human-rights due diligence. This often includes external human-rights assessments, which identify risks that we review carefully and decide how to address. We maintain a clear position on requests from governments for access to data. We also recently announced the Trusted Cloud Principles initiative, led by Google, Amazon, Microsoft, and other technology companies, to protect the rights of customers as they move to the cloud.From the beginning, Google’s mission has been to organize the world’s information and make it universally accessible and useful. Within Google Cloud, we aim to do the same for enterprise organizations, in ways that meet international and local standards. As the global landscape continues to evolve, we are committed to collaborating with human rights organizations and the broader technology industry to uphold human rights in every country where we operate. Learn more about our human rights efforts and global cloud infrastructure.Related ArticleExpanding our global footprint with new cloud regionsGoogle Cloud is expanding its global network with new regions in South America, Europe, and Asia.Read Article
Quelle: Google Cloud Platform

Google showcases Cloud TPU v4 Pods for large model training

Recently, models with billions or trillions of parameters have shown significant advances in machine learning capabilities and accuracy. For example, Google’s LaMDA model is able to engage in a free-flowing conversation with users about a large variety of topics. There is enormous interest within the machine learning research and product communities in leveraging large models to deliver breakthrough capabilities. The high computational demand of these large models requires an increased focus on improving the efficiency of the model training process, and benchmarking is an important means to coalesce the ML systems community towards realizing higher efficiencies.In the recently concluded MLPerf v1.1 Training round1, Google submitted two large language model benchmarks into the Open division, one with 480 billion parameters and a second with 200 billion parameters. These submissions make use of publicly available infrastructure, including Cloud TPU v4 Pod slices and the Lingvo open source modeling framework. Traditionally, training models at these scales would require building a supercomputer at a cost of tens or even hundreds of millions of dollars – something only a few companies can afford to do. Customers can achieve the same results using exaflop-scale Cloud TPU v4 Pods without incurring the costs of installing and maintaining an on-premise system. Large model benchmarksGoogle’s Open division submissions consist of a 480 billion parameter dense Transformer-based encoder-only benchmark using TensorFlow and a 200 billion-parameter JAX benchmark. These models are architecturally similar to MLPerf’s BERT model but with larger dimensions and number of layers. These submissions demonstrate large model scalability and high performance on TPUs across two distinct frameworks. Notably, these benchmarks, with their stacked transformer architecture, are fairly comparable in terms of their compute characteristics with other large language models.Figure 1: Architecture of the Encoder-only model used in Google’s MLPerf 1.1 submissions.Our two submissions were benchmarked on 2048-chip and 1024-chip TPU v4 Pod slices, respectively. We were able to achieve an end-to-end training time of ~55 hours for the 480B parameter model and ~40 hours for the 200B parameter model. Each of these runs achieved a computational efficiency of 63%- calculated as a fraction of floating point operations of the model together with compiler rematerialization over the peak FLOPs of the system used. Next-generation ML infrastructure for large Model training Achieving these impressive results required a combination of several cutting edge technologies. First, each TPU v4 chip provides more than 2X the compute power of a TPU v3 chip – up to 275 peak TFLOPS. Second, 4,096 TPU v4 chips are networked together into a Cloud TPU v4 Pod by an ultra-fast interconnect that provides 10x the bandwidth per chip at scale compared to typical GPU-based large scale training systems. Large models are very communication intensive: local computation often depends on results from remote computation that are communicated across the network. TPU v4’s ultra-fast interconnect has an outsized impact on computational efficiency of large models by eliminating latency and congestion in the network.Figure 2: A portion of one of Google’s Cloud TPU v4 Pods, each of which is capable of delivering in excess of 1 exaflop/s of computing power.The performance numbers demonstrated by our submission also rely on our XLA linear algebra compiler and leverage the Lingvo framework. XLA transparently performs a number of optimizations, including GSPMD based automatic parallelization of many of the computation graphs that form the building blocks of the ML model. XLA also allows for reduction in latency by overlapping communication with the computations. Our two submissions demonstrate the versatility and performance of our software stack across two frameworks, TensorFlow and JAX.Large models in MLPerfGoogle’s submissions represent an important class of models that have become increasingly important in ML research and production, but are currently not represented in MLPerf’s Closed division benchmark suite. We believe that adding these models to the benchmark suite is an important next step and can inspire the ML systems community to focus on addressing the scalability challenges that large models present.Our submissions demonstrate 63% computational efficiency, cutting edge in the industry. This high computational efficiency enables higher experimentation velocity through faster training. This directly translates into cost savings for Google’s Cloud TPU customers.  Please visit the Cloud TPU homepage and documentation to learn more about leveraging Cloud TPUs using TensorFlow, PyTorch, and JAX.1. The MLPerf name and logo are trademarks of MLCommons Association in the United States and other countries. All rights reserved. Unauthorized use is strictly prohibited. See www.mlcommons.org for more information.
Quelle: Google Cloud Platform

Docker SSO is Coming

The impending winter and holiday season hasn’t slowed us down here at Docker HQ. In fact, our engineers have been hard at work to put the finishing touches on one of our most requested features by our enterprise customers: Docker Single Sign-On (SSO).

With Docker SSO enabled, users can authenticate using their organization’s standard identity provider (IdP). This makes it easier for new users to quickly get started with Docker using their organization-provided email and existing password and also helps large organizations scale their use of Docker in a more manageable and secure way. To further simplify implementation, Docker works with a number of popular SAML IdPs including Google, Okta, Azure Active Directory, and more. Docker SSO is exclusive to Docker Business subscribers, and it is not included with the other Docker subscription tiers.

Now for the best part

We’re now welcoming a few of our current Docker Business customers to preview Docker SSO before it is generally available in January 2022. By giving some of our customers early access we hope to collect valuable feedback and data to ensure a seamless experience for all our users. 

If you currently have a Docker Business subscription and would like to preview Docker SSO for your organization, please let us know. We will contact you with instructions if you meet our eligibility criteria for early access. Not a Docker Business customer? Consider making the move today for access to Docker SSO and other premier features for management and security at scale.

We hope you are as excited about this upcoming Docker Business release as we are. Stay tuned for more.

DockerCon Live 2022  

Join us for DockerCon Live 2022 on Tuesday, May 10. DockerCon Live is a free, one day virtual event that is a unique experience for developers and development teams who are building the next generation of modern applications. If you want to learn about how to go from code to cloud fast and how to solve your development challenges, DockerCon Live 2022 offers engaging live content to help you build, share and run your applications. Register today at https://www.docker.com/dockercon/
The post Docker SSO is Coming appeared first on Docker Blog.
Quelle: https://blog.docker.com/feed/