BigLake: unifying data lakes and data warehouses across clouds

The volume of valuable data that organizations have to manage and analyze is growing at an incredible rate. This data is increasingly distributed across many locations, including  data warehouses, data lakes, and NoSQL stores. As an organization’s data gets more complex and proliferates across disparate data environments, silos emerge, creating increased risk and cost, especially when that data needs to be moved. Our customers have made it clear; they need help. That’s why today, we’re excited to announce BigLake, a storage engine that allows you to unify data warehouses and lakes. BigLake gives teams the power to analyze data without worrying about the underlying storage format or system, and eliminates the need to duplicate or move data, reducing cost and inefficiencies. With BigLake, users gain fine-grained access controls, along with performance acceleration across BigQuery and multicloud data lakes on AWS and Azure. BigLake also makes that data uniformly accessible across Google Cloud and open source engines with consistent security. BigLake extends a decade of innovations with BigQuery to data lakes on multicloud storage, with open formats to ensure a unified, flexible, and cost-effective lakehouse architecture.BigLake architectureBigLake enables you to:Extend BigQuery to multicloud data lakes and open formats such as Parquet and ORC with fine-grained security controls, without needing to set up new infrastructure.Keep a single copy of data and enforce consistent access controls across analytics engines of your choice, including Google Cloud and open-source technologies such as Spark, Presto, Trino, and Tensorflow.Achieve unified governance and management at scale through seamless integration with Dataplex.Bol.com, an early customer using BigLake, has been accelerating analytical outcomes while keeping their costs low:“As a rapidly growing e-commerce company, we have seen rapid growth in data. BigLake allows us to unlock the value of data lakes by enabling access control on our views while providing a unified interface to our users and keeping data storage costs low. This in turn allows quicker analysis on our datasets by our users.”—Martin Cekodhima, Software Engineer, Bol.comExtend BigQuery to unify data warehouses and lakes with governance across multicloud environmentsBy creating BigLake tables, BigQuery customers can extend their workloads to data lakes built on Google Cloud Storage (GCS), Amazon S3, and Azure data lake storage Gen 2. BigLake tables are created using a cloud resource connection, which is a service identity wrapper that enables governance capabilities. This allows administrators to manage access control for these tables similar to BigQuery tables, and removes the need to provide object store access to end users. Data administrators can configure security at the table, row or column level on BigLake tables using policy tags. For BigLake tables defined over Google Cloud Storage, fine grained security is consistently enforced across Google Cloud and supported open-source engines using BigLake connectors. For BigLake tables defined on Amazon S3 and Azure data lake storage Gen 2, BigQuery Omni enables governed multicloud analytics by enforcing security controls. This enables you to manage a single copy of data that spans BigQuery and data lakes, and creates interoperability between data warehousing, data lake, and data science use cases.Open interface to work consistently across analytic runtimes spanning Google Cloud technologies and open source engines Customers running open source engines like Spark, Presto, Trino, and Tensorflow through Dataproc or self managed deployments can now enable fine-grained access control over data lakes, and accelerate the performance of their queries. This helps you build secure and governed data lakes, and eliminate the need to create multiple views to serve different user groups. This can be done by creating BigLake tables from a supported query engine like Spark DDL, and using Dataplex to configure access policies. These access policies are then enforced consistently across the query engines that access this data – greatly simplifying access control management. Achieve unified governance & management at scale through seamless integration with DataplexBigLake integrates with Dataplex to provide management-at-scale capabilities. Customers can logically organize data from BigQuery and GCS into lakes and zones that map to their data domains, and can centrally manage policies for governing that data. These policies are then uniformly enforced by Google Cloud and OSS query engines. Dataplex also makes management easier by automatically scanning Google Cloud storage to register BigLake table definitions in BigQuery, and makes them available via Dataproc Metastore. This helps end users discover these BigLake tables for exploration and querying using both OSS applications and BigQuery. Taken together, these capabilities enable you to run multiple analytic runtimes over data spanning lakes and warehouses in a governed manner. This breaks down data silos and significantly reduces the infrastructure management, helping you to advance your analytics stack and unlock new use cases.What’s next?If you would like to learn more about BigLake, please visit our website. Alternatively, get started with BigLake today by using this quickstart guide, or contact the Google Cloud sales team.Related ArticleLimitless Data. All Workloads. For EveryoneRead about the newest innovations in data cloud announced at Google Cloud’s Data Cloud Summit.Read Article
Quelle: Google Cloud Platform

Bringing together the best of both sides of BI with Looker and Data Studio

In today’s data world, companies have to consider many scenarios in order to deliver insights to users in a way that makes sense to them. On one hand, they must deliver trusted, governed reporting to inform mission-critical business functions. But on the other hand, they must enable a broad population to answer questions on their own, in an agile and self-serve manner. Data-driven companies want to democratize access to relevant data and empower their business users to get answers quickly. The trade-off is governance; it can be hard to maintain a single shared source of truth, granular control of data access, and secure centralized storage. Speed and convenience compete with trust and security. Companies are searching for a BI solution that can be easily used by everyone in an organization (both technical and non-technical), while still producing real-time governed metrics.Today, a unified experience for both self-serve and governed BI gets one step closer, with our announcement of the first milestone in our integration journey between Looker and Data Studio. Users will now be able to access and import governed data from Looker within the Data Studio interface, and build visualizations and dashboards for further self-serve analysis.This union allows us to better deliver a complete, integrated Google Cloud BI experience for our users. Business users will feel more empowered than ever, while data leaders will be able to preserve the trust and security their organization needs.This combination of ​self-serve and governed BI together will help enterprises make better data-driven decisions. Looker and Data Studio, Better Together Looker is a modern enterprise platform for business intelligence and analytics, that helps organizations build data-rich experiences tailored to every part of their business. Data Studio is an easy to use self-serve BI solution enabling ad-hoc reporting, analysis, and data mashups across 500+ data sets. Looker and Data Studio serve complementary use cases. Bringing Looker and Data Studio together opens up exciting opportunities to combine the strengths of both products and a broad range of BI and analytic capabilities to help customers reimagine the way they work with data. How will the integration work?This first integration between these products allows users to connect to the Looker semantic model directly from Data Studio. Users can bring in the data they wish to analyze, connect it to other available data sources, and easily explore and build visualizations within Data Studio. The integration will follow three principles with respect to governance:Access to data is enforced by Looker’s security features, the same way it is when using the Looker user interface.Looker will continue to be the single access point for your data.Administrators will have full capabilities to manage Data Studio and Looker together.How are we integrating these products?This is the first step in our roadmap that will bring these two products closer together. Looker’s semantic model allows metrics to be centrally defined and broadly used, ensuring a single version of the truth across all of your data. This integration allows the Looker semantic model to be used within Data Studio reports, allowing people to use the same tool to create reports that rely on both ad-hoc and governed data. This brings together the best of both worlds – a governed data layer, and a self-serve solution that allows analysis of both governed and ungoverned data.With this announcement, the following use cases will be supported:Users can turn their Looker-governed data into informative, highly customizable dashboards and reports in Data Studio.Users can blend governed data from Looker with data available from over 500 data sources in Data Studio, to rapidly generate new insights.Users can analyze and rapidly prototype ungoverned data (from spreadsheets, csv files, or other cloud sources) within Data Studio.Users can collaborate in real-time to build dashboards with teammates or people outside the company. When will this integration be available to use?The Data Studio connector for Looker is currently in preview. If you are interested in trying it out, please fill out this form.Next StepsThis integration is the first of many in our effort to bring Looker and Data Studio closer together. Future releases will introduce additional features to create a more seamless user experience across these two products. We are very excited to roll out new capabilities in the coming months and will keep you updated on our future integrations of the two products.Related ArticleLimitless Data. All Workloads. For EveryoneRead about the newest innovations in data cloud announced at Google Cloud’s Data Cloud Summit.Read Article
Quelle: Google Cloud Platform

The future is on FHIR for SAS and Microsoft Azure

This blog has been co-authored by Steve Kearney, PharmD, Global Medical Director, SAS.

This blog is part of a series in collaboration with our partners and customers leveraging the newly announced Azure Health Data Services. Azure Health Data Services, a platform as a service (PaaS) offering designed to support Protected Health Information (PHI) in the cloud, is a new way of working with unified data—providing care teams with a platform to support both transactional and analytical workloads from the same data store and enabling cloud computing to transform how we develop and deliver AI across the healthcare ecosystem.

There is a dichotomy in health care technology. Despite new developments in imaging, diagnostics, treatment, and surgical techniques, the lack of data standardization in the industry has trapped health insights in functional silos. Providers and payers alike struggle to manually reconcile incompatible file formats, which slows the transfer of information and negatively impacts quality care and patient experience.

Microsoft, along with partners such as global analytics software company SAS, are driving towards increased interoperability through enabling the use of standards such as Fast Healthcare Interoperability Resources (FHIR®). Together, SAS and Microsoft Azure are building deep technology integrations that unlock value by making disparate data and advanced analytics more accessible to health and life science organizations. With new capabilities such as the integration from Azure Health Data Services to SAS on Azure, the embedded AI capabilities of SAS Health are more efficient and secure, expanding the possibilities of patient-centric innovation and trusted collaboration across the health landscape.

FHIR puts the patient at the center of the health care ecosystem. When querying information in the previous HL7 format, the query is answered with the entire patient dataset that must be parsed to find the information desired for predictive modeling. Additionally, data would require harmonization within and across the organization, creating limitations on available data. In contrast, harmonized FHIR datasets persisting on Azure Health Data Services enable FHIR-based requests directed to the specific data points required, speeding up queries to near-real-time and protecting patient data.

While FHIR’s footprint in the industry is small compared to HL7’s, the global adoption of the FHIR standard is growing. Major electronic health records (EHR) companies like Cerner and Epic are moving quickly to support FHIR.1 Notably in the United States, the Centers for Medicare and Medicaid Services (CMS) has mandated its use for health insurance payers and providers.

Transform your analytical experience in the health cloud

The integration between Azure Health Data Services and SAS Health can be transformational for organizations who have struggled to operationalize analytics. Not only does this integration offer a technology that is secure, fast, and scalable, it democratizes analytics by allowing the business or clinical user to query a patient data set using a pre-set parameter or algorithm and return results within a clinical workflow.

The traditional view of health analytics is that it occurs outside the process of care and is in some way removed from the patient. That’s changing, thanks to secure health cloud environments like Azure Health Data Services and presents the opportunity for more real-time integration of patient and claims data. With the evolution of the citizen data scientist and respective interoperability, we now see a clearer path from analytics to improved health care outcomes.

The graphic below illustrates the role of health data analytic interoperability in health and life sciences. Ultimately, the use of diverse health data throughout the process of care in a shared cloud environment will enable better outcomes for us all.

SAS Health and Azure Health Data Services

The embedded-AI capabilities of SAS Health running on FHIR data ingested through Azure Health Data Services provide game-changing advantages across health care delivery and research.

Providers

SAS Health on FHIR gives speedy access to analytic insights within EHRs, parsing out only the information needed, allowing near-real-time results from, for example, pharmacy claims, laboratory results, or imaging. Predictive insights such as medication adherence or emerging health risks are more available through a secure FHIR-based exchange. Quality care and patient satisfaction increase when providers can integrate data across multiple systems and record types including patient records and claims data into a single view.

Payers

Payers governed by CMS are already mandated to transition to FHIR-based communication standards and are experiencing early wins. For example, adjudication of claims is one of the most time-consuming parts of the payer process. With FHIR, payers can securely query patient records to determine medical necessity of a service or procedure and whether appropriate authorization was obtained, cutting time dramatically in the process. With FHIR’s extensibility beyond the payer-provider core, pharmacy data can be queried to inform proactive disease management programs with specialty drugs and more real-time formulary approvals to meet patient needs.

Academic researchers

For clinical research, data sharing can be a common, time-consuming obstacle. FHIR-ready datasets can accelerate the generation of new health insights and expand the universe of data types for research, including social determinants of health, real-world data, genetics, device data from the internet of medical things, and more.

Ultimately, these innovations in health data analytic interoperability can make insights faster across the vast ecosystem of professionals who are committed to a healthier world. While technology is only one part of the solution, improving health begins with predicting future health risks and taking proactive steps to mitigate disease and promote physical and mental wellness.

Do more with your data with Microsoft Cloud for Healthcare

With Azure Health Data Services, health organizations can transform their patient experience, discover new insights with the power of machine learning and AI, and manage PHI data with confidence. Enable your data for the future of healthcare innovation with Microsoft Cloud for Healthcare.

We look forward to being your partner as you build the future of health.

Learn more about Azure Health Data Services.
Learn more about SAS Health on Azure.
Read our recent blog, “Microsoft launches Azure Health Data Services to unify health data and power AI in the cloud.”
Learn more about Microsoft Cloud for Healthcare.

®FHIR is a registered trademark of Health Level Seven International, registered in the U.S. Trademark Office and are used with their permission.

1Journal of the American Medical Informatics Association, Volume 28, Issue 11, November 2021, pages 2379–2384.
Quelle: Azure

Azure delivers strong MLPerf inferencing v2.0 results from 1 to 8 GPUs

Microsoft Azure is committed to providing its customers with industry-leading real-world AI capabilities. In December 2021, Microsoft Azure debuted its leadership performance with the MLPerf training v1.1 results. Azure debuted at number one among cloud providers and number two overall at scale among all submitters. Azure’s supercomputer's building blocks were used to generate the results in our v2.0 submissions for the MLPerf inferencing results published on April 6, 2022.

These industry-leading results are driven by Microsoft’s publicly available supercomputing capabilities designed for real-world AI inferencing workloads. Microsoft enables customers of all scales to deploy powerful AI solutions, whether at a focused local scale or at the scale of the largest supercomputers in the world.

Microsoft Azure’s publicly available AI inferencing capabilities are led by the NDm A100 v4, ND A100 v4, and NC A100 v4 virtual machines (VMs) that are powered by NVIDIA A100 SXM and PCIe Tensor Core graphics processing units (GPUs). These results showcase Azure’s commitment to making AI inferencing available to all in the most accessible way—while raising the bar for AI inferencing in Azure.

In our quest to continually provide the best technology for our customers, Azure has recently announced the preview for the NC A100 v4. With this introduction of the NC A100 v4 series, we have provided our customers with three different VM sizes ranging from one to four GPUs. From our benchmarking, we have seen more than two times performance over the previous generation. Azure’s customers can get access to these new systems today by signing up for the preview program.

Some highlights for this round of MLPerf inferencing submissions can be seen in the following tables.

Highlights from the results

ND96amsr A100 v4 powered by NVIDIA A100 80G SXM Tensor Core GPU

Benchmark
Samples/second
Queries/second
Scenarios

bert-99
27,500 plus
~22,500 plus
Offline and server

resnet
300,000 plus
~200,000 plus
Offline and server

3d-unet
24.87
 
Offline

NC96ads A100 v4 powered by NVIDIA A100 80G PCIe Tensor Core GPU

Benchmark
Samples/second
Queries/second
Scenarios

bert-99
~6,300
~5,300
Offline and server

resnet
144,000
~119,600
Offline and server

3d-unet
11.7
 
Offline

The above tables showcase three of the six benchmarks the team ran using NVIDIA A100 SXM and PCIe Tensor Core GPUs for offline and server scenarios respectively. Take a look at the full list of results for the various divisions.

Azure works closely with NVIDIA

The results were generated by deploying the environment using the VM offerings and Azure’s Ubuntu 18.04-HPC marketplace image. We worked closely with NVIDIA to quickly deploy the environment and perform benchmarks with industry-leading results in performance and scalability.

These results are a testament to Azure’s focus on offering scalable supercomputing for any workload while enabling our customers to utilize “on-demand” supercomputing capabilities in the cloud to solve their most complex problems. Visit the Azure Tech Community blog to read the steps to reproduce the results.

More about MLPerf

MLPerf is a consortium of AI leaders from academia, research labs, and industry where the mission is to “build fair and useful benchmarks” that provide unbiased evaluations of training and inference performance for hardware, software, and services—all conducted under prescribed conditions. To stay on the cutting edge of industry trends, MLPerf continues to evolve, holding new tests at regular intervals and adding new workloads that represent state-of-the-art AI. MLPerf’s tests are transparent and objective, so users can rely on the results to make informed buying decisions. The industry benchmarking group, formed in May 2018, is backed by dozens of industry leaders. The benchmark tests across inferencing are increasingly becoming the key tests that hardware and software vendors use to demonstrate performance. Take a look at the full list of results for MLPerf Inference v2.0.
Quelle: Azure

Limitless Data. All Workloads. For Everyone

Today, data exists in many formats, is provided in real-time streams, and stretches across many different data centers and clouds, all over the world. From analytics, to data engineering, to AI/ML, to data-driven applications, the ways in which we leverage and share data continues to expand. Data has moved beyond the analyst and now impacts every employee, every customer, and every partner. With the dramatic growth in the amount and types of data, workloads, and users, we are at a tipping point where traditional data architectures – even when deployed in the cloud – are unable to unlock its full potential. As a result, the data-to-value gap is growing. To address these challenges, we are unveiling several data cloud innovations today that allow our customers to work with limitless data, across all workloads, and extend access to everyone. These announcements include BigLake and Spanner change streams to further unify customer data while ensuring it’s delivered in real-time, as well as Vertex AI Workbench and Model Registry to close the data to AI value gap. And to bring data within reach for anyone, we are announcing a unified business intelligence (BI) experience that includes a new Workspace integration, along with new programs that further enable our data cloud partner ecosystem. Removing all data limits Today, we are announcing the preview of BigLake, a data lake storage engine, to remove data limits by unifying data lakes and warehouses. Managing data across disparate lakes and warehouses creates silos and increases risk and cost, especially when data needs to be moved. BigLake allows companies to unify their data warehouses and lakes to analyze data without worrying about the underlying storage format or system, which eliminates the need to duplicate or move data from a source and reduces cost and inefficiencies. With BigLake, customers gain fine-grained access controls, with an API interface spanning Google Cloud and open file formats like Parquet, along with open-source processing engines like Apache Spark. These capabilities extend a decade’s worth of innovations with BigQuery to data lakes on Google Cloud Storage to enable a flexible and cost-effective open lake house architecture. Twitter already uses storage capabilities with BigQuery to remove the limits of data to better understand how people use their platform, and what types of content they might be interested in. As a result, they are able to serve content across trillions of events per day with an ads pipeline that runs more than 3M aggregations per second. Another major innovation we’re announcing today is Spanner change streams. Coming soon, this new product will further remove data limits for our customers, allowing them to track changes within their Spanner database in real time in order to unlock new value. Spanner change streams tracks Spanner inserts, updates, and deletes to stream the changes in real time across a customer’s entire Spanner database. This ensures customers always have access to the freshest data as they can easily replicate changes from Spanner to BigQuery for real-time analytics, trigger downstream application behavior using Pub/Sub, or store changes in Google Cloud Storage (GCS) for compliance. With the addition of change streams, Spanner, which currently processes over 2 billion requests per second at peak with up to 99.999% availability, now gives customers endless possibilities to process their data. Remove the limits of your data workloadsOur AI portfolio is powered by Vertex AI, a managed platform with every ML tool needed to build, deploy and scale models, and is optimized to work seamlessly with data workloads in BigQuery and beyond. Today, we’re announcing new Vertex AI innovations that will provide customers with an even more streamlined experience to get AI models into production faster and make maintenance even easier.Vertex AI Workbench, which is now generally available, brings data and ML systems into a single interface so that teams have a common toolset across data analytics, data science, and machine learning. With native integrations across BigQuery, Serverless Spark, and Dataproc, Vertex AI Workbench enables teams to build, train and deploy ML models 5X faster than traditional notebooks. In fact, a global retailer was able to drive millions of dollars in incremental sales and deliver 15% faster speed to market with Vertex AI Workbench.With Vertex AI, customers have the ability to regularly update their models. But managing the sheer number of artifacts involved can quickly get out of hand. To make it easier to manage the overhead of model maintenance, we are announcing new MLOps capabilities with Vertex AI Model Registry. Now in preview, Vertex AI Model Registry provides a central repository for discovering, using, and governing machine learning models, including those in BigQuery ML. This makes it easy for data scientists to share models and application developers to use them, ultimately enabling teams to turn data into real-time decisions, and be more agile in the face of shifting market dynamics.Extending the reach of your dataToday, we are launching Connected Sheets for Looker, and the ability to access Looker data models within Data Studio. Customers now have the ability to interact with data however they choose, whether it be through Looker Explore, from Google Sheets, or using the drag-and-drop Data Studio interface. This will make it easier for everyone to access and unlock insights from data in order to drive innovation, and to make data-driven decisions with this new unified Google Cloud business intelligence (BI) platform. This unified BI experience makes it easy to tap into governed, trusted enterprise data, to incorporate new data sets and calculations, and to collaborate with peers.Mercado Libre, the largest online commerce and payments ecosystem in Latin America, has been an early adopter of Connected Sheets for Looker. Using this integration, they have been able to provide broader access to data through a spreadsheet interface that their employees are already familiar with. By lowering the barrier to entry, they have been able to build a data-driven culture in which everyone can inform their decisions with data. Doubling down on the data cloud partner ecosystemClosing the data-to-value gap with these data innovations would not be possible without our incredible partner ecosystem. Today, there are more than 700 software partners powering their applications using Google’s data cloud. Many partners like Bloomreach, Equifax, Exabeam, Quantum Metric, and ZoomInfo, have started using our data cloud capabilities with the Built with BigQuery initiative, which provides access to dedicated engineering teams, co-marketing, and go-to-market support. Our customers want partner solutions that are tightly integrated and optimized with products like BigQuery. So today, we’re announcing Google Cloud Ready – BigQuery, a new validation that recognizes partner solutions like those from Fivetran, Informatica and Tableau that meet a core set of functional and interoperability requirements. Today, we already recognize more than 25 partners in this new Google Cloud Ready – BigQuery program that reduces costs for customers associated with evaluating new tools while also adding support for new customer use cases. We’re also announcing a new Database Migration Program to help our customers efficiently and effectively accelerate the move from on-premise and other clouds to Google’s industry-leading managed database services. This includes tooling, resources, and knowledgeable experience from alliances like Deloitte, as well as incentives from Google to offset the cost of migrating databases.We remain committed to continued innovation with the leading data and analytics companies where our customers are investing. This week Databricks, Fivetran, MongoDB, Neo4j, and Redis are all announcing significant new capabilities for customers on Google Cloud.All of these announcements and more will be shared in detail at our Data Cloud Summit. Be sure to watchthe data cloud strategy sessions, breakouts, and get access to hands on content. There is no doubt the future of data holds limitless possibilities, and we are thrilled to be on this data cloud journey.Related ArticleReady to solve for the future? Data Cloud Summit ’22 is coming April 6Hear from customers, leaders and builders from Google Cloud at Data Cloud Summit 2022 to get the insight you need for your data organizationRead Article
Quelle: Google Cloud Platform