Getting Started with Google Cloud Logging Python v3.0.0

We’re excited to announce the release of a major update to the Google Cloud Python logging library. v3.0.0 makes it even easier for Python developers to send and read logs from Google Cloud, providing real-time insights into what is happening in your application.  If you’re a Python developer working with Google Cloud, now is a great time to try out Cloud Logging!If you’re unfamiliar with the `google-cloud-logging` library, getting started is simple. First, download the library using pip:Now, you can set up the client library to work with Python’s built-in `logging` library. Doing this will make it so that all your standard Python log statements will start sending data to Google Cloud:We recommend using the standard Python `logging` interface for log creation, as demonstrated above. However, if you need access to other Google Cloud Logging features (reading logs, managing log sinks, etc), you can use `google.cloud.logging` directly:Here are some of the main features of the new release:Support More Cloud EnvironmentsPrevious versions of google-cloud-logging supported onlyApp Engine andKubernetes Engine. Users reported that the library would occasionally drop logs on serverless environments like Cloud Run and Cloud Functions. This was because the library would send logs in batches over the network. When a serverless environment would spin down, unsent batches could be lost.v3.0.0 fixes this issue by making use of GCP’s built instructured JSON logging functionality on supported environments (GKE, Cloud Run, or Cloud Functions). If the library detects it is running on an environment that supports structured logging, it will automatically make use of the newStructuredLogHandler, which writes logs as JSON strings printed to standard out. Google Cloud’s built-in agents will then parse the logs and deliver them to Cloud Logging, even if the code that produced the logs has spun down. Structured Logging is more reliable on serverless environments, and it allows us to support all major GCP compute environments in v3.0.0. Still, if you would prefer to send logs over the network as before, you can manually set up the library with a CloudLoggingHandler instance:Metadata AutodetectionWhen you troubleshoot your application, it can be useful to have as much information about the environment as possible captured in your application logs. `google-cloud-logging` attempts to help in this process by detecting and attaching metadata about your environment to each log message. The following fields are currently supported:`resource`: The Google Cloud resource the log originated from for example, Functions, GKE, or Cloud Run`httpRequest`: Information about an HTTP request in the log’s contextFlask and Django are currently supported`sourceLocation` : File, line, and function namestrace, spanId, and traceSampled: Cloud Trace metadataSupports X-Cloud-Trace-Context and w3c transparent trace formatsThe library will make an attempt to populate this data whenever possible, but any of these fields can also be explicitly set by developers using the library.JSON Support in Standard Library IntegrationGoogle Cloud Logging supports bothstring and JSON payloads for LogEntries, but up until now,the Python standard library integration could only send logs with string payloads.In `google-cloud-logging` v3,  you can log JSON data in two ways:1. Log a JSON-parsable string:2. Pass a `json_fields` dictionary using Python logging’s `extra` argument:Next StepsWith version v3.0.0, the Google Cloud Logging Python library now supports more compute environments, detects more helpful metadata, and provides more thorough support for JSON logs. Along with these major features, there are also user-experience improvements like a new log method and more permissive argument parsing. If you want to learn more about the latest release, these changes and others are described in more detail in the v3.0.0 Migration Guide. If you’re new to the library, check out the google-cloud-logging user guide. If you want to learn more about observability on GCP in general, you can spin up test environments using Cloud Ops Sandbox.Finally, if you have any feedback about the latest release, have new feature requests, or would like to make any contributions, feel free to open issues on our GitHub repo. The Google Cloud Logging libraries are open source software, and we welcome new contributors!Related ArticleTake the first step toward SRE with Cloud Operations SandboxSpin up the Cloud Operations Sandbox to see how Google’s logging, monitoring, tracing, profiling and debugging can kickstart your SRE pra…Read Article
Quelle: Google Cloud Platform

Genomic analysis on Galaxy using Azure CycleCloud

Cloud computing and digital transformation have been powerful enablers for genomics. Genomics is expected to be an exabase-scale big data domain by 2025, posing data acquisition and storage challenges on par with other major generators of big data. Embracing digital transformation offers a practically limitless ability to meet the genomic science demands in both research and medical institutions. The emergence of cloud-based computing platforms such as Microsoft Azure has paved the path for online, scalable, cost-effective, secure, and shareable big data persistence and analysis with a growing number of researchers and laboratories hosting (publicly and privately) their genomic big data on cloud-based services.

At Microsoft, we recognize the challenges faced by the genomics community and are striving to build an ecosystem (backed by OSS and Microsoft products and services) that can facilitate genomics work for all. We’ve focused our efforts on three main core areas—research and discovery in genomic data, building out a platform to enable rapid automation and analysis at scale, and optimized and secure pipelines at a clinical level. One of the core Azure services that has enabled us to leverage high performance compute environment to perform genomic analysis is Azure CycleCloud.

Galaxy and Azure CycleCloud

Galaxy is a scientific workflow, data integration, and data analysis persistence and publishing platform that aims to make computational biology accessible to research scientists that do not have computer programming or systems administration experience. Although it was initially developed for genomic research, it is largely domain agnostic and is now used as a general bioinformatics workflow management system. Galaxy system is used for accessible, reproducible, and transparent computational research.

Accessible: Programming experience is not required to easily upload data, run complex tools and workflows, and visualize results.
Reproducible: Galaxy captures information so that you don't have to; any user can repeat and understand a complete computational analysis, from tool parameters to the dependency tree.
Transparent: Users share and publish their histories, workflows, and visualizations via the web.
Community-centered: Inclusive and diverse users (developers, educators, researchers, clinicians, and more) are empowered to share their findings.

Azure CycleCloud is an enterprise-friendly tool for orchestrating and managing high-performance computing (HPC) environments on Azure. With Azure CycleCloud, users can provision infrastructure for HPC systems, deploy familiar HPC schedulers, and automatically scale the infrastructure to run jobs efficiently at any scale. Through Azure CycleCloud, users can create different types of file systems and mount them to the compute cluster nodes to support HPC workloads. With dynamic scaling of clusters, the business can get the resources it needs at the right time and the right price. Azure CycleCloud automated configuration enables IT to focus on providing service to the business users.

Deploying Galaxy on Azure using Azure CycleCloud

Galaxy is used by most academic institutions that conduct genomic research. Most institutions that already use Galaxy want to stick to it because it provides multiple tools for genomic analysis as a SaaS platform. Users can also deploy custom tools onto Galaxy.

Galaxy users generally use the SaaS version of Galaxy as part of UseGalaxy resources. UseGalaxy servers implement a common core set of tools and reference genomes and are open to anyone to use. All information on its usage is available on the Galaxy Platform Directory.

However, there are some research institutions that intend to deploy Galaxy in-house as an on-premises solution or a cloud-based solution. The remainder of this article describes how to deploy and run Galaxy on Microsoft Azure using Azure CycleCloud and grid engine cluster. The solution was built during the Microsoft hackathon (October 12 to 14, 2021) with code implementation assistance from Azure HPC Specialist, Jerry Morey. The architectural pattern described below can help organizations to deploy Galaxy in an Azure environment using CycleCloud and a scheduler of choice.

As a pre-requisite, genomic data should be available in a storage location, either cloud or on-premises. Azure CycleCloud should be deployed using the steps described in the “Install CycleCloud using the Marketplace image” documentation.

Cluster deployment that is truly supported by Galaxy on the cloud is called the unified method. In this method, the copy of Galaxy on the application server is the same copy as the one on the cluster nodes. The most common method to do this would be to put Galaxy in a network file system (NFS) somewhere that is accessible by the application server and the cluster nodes. This is the most common deployment method for Galaxy.

An admin user can SSH into Azure CycleCloud virtual machines or Galaxy server virtual machines to perform admin-related activities. It is recommended to close the SSH port when in production. Once the Galaxy server is running on a node, end users (researchers) can load the portal on their end device to perform analysis tasks which include loading data, installing, uploading tools, and more.

Access to functionalities (such as installing and deleting tools versus the usage of tools for analysis) are controlled by parameters defined in galaxy.yml that resides in the Galaxy server. Once a user accesses a functionality, they are converted to jobs that are submitted to the grid engine cluster for further execution.

Deployment scripts are available to ease deployment. These scripts can be used to deploy the latest version of Galaxy on Azure CycleCloud.
Following are the steps to use the deployment scripts:

Git clone this project (The project is in active development, so cloning the latest release is recommended).

git clone –b release_21.09 https://github.com/themorey/galaxy-gridengine.git

Upload project to CC locker.

cd galaxy-gridengine

Modify files (if needed)

cyclecloud locker list

Azure cycle Locker (az://mystorageaccount/cyclecloud

cyclecloud project upload "Azure cycle Locker"

Import cluster template to CC.

cyclecloud import_cluster <cluster-name> -c <galaxy-folder-name> -f templates/gridengine-galaxy2.txt

NOTE: Substitute <cluster-name> with a name for your cluster—all lower case, no spaces.

Navigate to CC Portal to configure and start the cluster.

Wait for 30 to 45 minutes for the Galaxy server to be installed.

To check if the server is installed correctly, SSH into Galaxy server node and check galaxy.log in /shared/home/<galaxy-folder-name> directory.

This deployment was adopted by a leading United States-based academic medical center. The Microsoft Industry Solutions team helped deploy this solution on the customer’s Azure tenant. Researchers at the center tested to assess the parity of this solution to existing Galaxy deployment on their on-premises HPC environment. They were able to successfully test the deployed Galaxy server that used Azure CycleCloud for job orchestration. Several common bioinformatics tools such as bedtools, fastqc, bcftools, picard, and snpeff were installed and tested. Galaxy supports local user by default. As part of this engagement, a solution to integrate their corporate active directory was tested and deployed. The solution was found to be on par with their on-premises deployment. With the increased number of execute nodes and size of those nodes, they found that the jobs were executed in less time.

For more information, support, or guidance related to the content in this blog, we recommend you reach out to your Microsoft sales representative.

Learn more

Learn more about Microsoft Genomics solutions.

Microsoft Genomics service on Azure.
Azure CycleCloud—HPC Cluster and Workload Management.
Galaxy on Azure deployment scripts.

Quelle: Azure

Learn how open source plays a key role in Microsoft’s cloud strategy with Inside Azure for IT

With more than 1 million views of our fireside chats, we’re inspired by the tremendous opportunity to connect those within the community—customers, partners, and technology enthusiasts everywhere. Whether you engage in the live ask-the-experts sessions, watch the deep-dive skilling videos, or join us for fireside chats—the Azure team and I are delighted and humbled by your participation and enthusiasm for Inside Azure for IT. 

In our third episode, we talk about some of our Linux and open source-related partnerships, product innovation, and initiatives, plus how that helps customers and communities. To those who think of Azure as “mostly Windows cloud,” it may be surprising to learn that more than 60 percent of Azure customer compute cores are Linux-based, and that Linux virtual machine (VM) cores are growing faster than those based on Windows.

My own career has mirrored Microsoft’s evolution of how we think about, contribute to, and consume Linux and open source. For example, I’ve gone from being solely focused on Windows and Windows Server, to learning how to contribute upstream to make Linux run great on Hyper-V, to now, where open source and Linux are core to the development of Azure.

In this episode, you’ll get a behind-the-scenes peek at Microsoft’s approach, and how we've brought together customers, partners, and communities to innovate and collaborate across open-source technologies.

Innovate with Open Source and Linux on Azure

The episode is divided into three separate segments so you can watch them individually on-demand at your convenience.

Part one: Microsoft and Red Hat on simplifying cloud adoption with joint innovation on Azure with Linux

In this segment, you’ll hear from Red Hat about partnering with Microsoft and how it helps customers with their cloud modernization and migration journey. Mike Evans, VP, Technical Business Development, and Xavier Lecauchois, Sr. Director Ansible Cloud Services, from Red Hat join me to chat about the strategy and the latest innovation, the Red Hat Ansible Automation Platform on Azure. Watch: Simplifying cloud adoption with joint innovation on Azure with Linux.

Part two: Brendan Burns and Krishna Ganugapati on safeguarding workloads with Mariner—Microsoft’s internal Linux distro

Delivering reliable Azure services to customers faster is the driving force behind the creation of Mariner, Microsoft’s own Linux distro. Join me, as I chat with Krishna Ganugapati, VP of Software Engineering, Edge OS, and Brendan Burns, CVP, Azure Cloud Native on why the Azure team created Mariner and how it’s benefiting customers and Microsoft engineers. Watch: Safeguarding workloads with Mariner—Microsoft Azure’s own internal Linux distro.

Part three: Microsoft's Sarah Novotny on working together with open source communities to drive innovation

Open source connects developers around the world, providing ways to collaborate and innovate collectively. Join Sarah Novotny, Director of Open-source Strategy for Azure as we chat about running open-source technologies in the cloud, how the relationship between IT and developers enables open-source innovation, Microsoft’s leadership and contributions to help secure open-source software, and her unique background in the open-source community. Watch: Developing in the open and working together to drive innovation.

Stay current with Inside Azure for IT

Beyond this latest episode, there are many more technical and cloud-skilling resources available through Inside Azure for IT. Learn more about empowering an adaptive IT environment with best practices and resources designed to enable productivity, digital transformation, and innovation. Take advantage of technical training videos and learn about implementing these scenarios.

Register for Azure Open Source Day to watch live on February 15, 2022, 9:00 AM to 10:30 AM Pacific Time or on-demand later.
Get started by learning about Linux on Azure.
See our schedule for Ask the product experts live.
Watch part one: Microsoft and Red Hat on simplifying cloud adoption with joint innovation on Azure with Linux.
Watch part two: Brendan Burns and Krishna Ganugapati on safeguarding workloads with Mariner—Microsoft Azure’s own internal Linux distro.
Watch part three: Microsoft’s Sarah Novotny on developing in the open and working together to drive innovation.

Quelle: Azure

AWS Secrets Manager unterstützt jetzt Drehungsfenster

AWS Secrets Manager unterstützt jetzt die Möglichkeit, geheime Drehungen innerhalb bestimmter Zeitfenster zu planen. Mit dieser Funktion können Sie Geheimnis-Drehungen auf bestimmte Stunden an bestimmten Tagen beschränken. Zuvor unterstützte Secrets Manager die automatische Drehung von Geheimnissen innerhalb der letzten 24 Stunden des angegebenen Drehungsintervalls. Mit der heutigen Einführung müssen Sie sich nicht mehr zwischen dem Komfort verwalteter Drehungen und der Betriebssicherheit von Wartungsfenstern entscheiden.
Quelle: aws.amazon.com

Amazon Comprehend launcht Modellkopie für benutzerdefiniertes Comprehend

Amazon Comprehend unterstützt jetzt die Modellkopierfunktion, die es Kunden ermöglicht, benutzerdefinierte Comprehend-Klassifizierungs- oder benutzerdefinierte Entitätserkennungsmodelle von einem AWS-Quellkonto in ein bestimmtes AWS-Zielkonto in derselben AWS-Region zu kopieren. Unternehmenskunden und AWS-Partner verwenden häufig mehrere AWS-Konten, die basierend auf Entwicklungsphasen (z. B. Entwicklung, Test, Staging, Bereitstellung), basierend auf einer Geschäftsfunktion (z. B. Datenwissenschaft, Engineering) oder einer Kombination aus beidem bereitgestellt werden. Bisher konnten benutzerdefinierte Comprehend-Modelle nur in dem AWS-Konto verwendet werden, in dem sie trainiert wurden. Dazu mussten Kunden ihren Trainingsdatensatz und ihre Anmerkungen in jedes AWS-Konto kopieren und ein separates Modell trainieren, was zeitaufwändig und teuer ist und die Bereitstellungsgeschwindigkeit insgesamt verringert.
Quelle: aws.amazon.com

Das Erhalten personalisierter Erlebnisse mit Machine Learning fügt Unterstützung für Geschäftsdomänen und Benutzersegmentierung hinzu

AWS Solutions hat die Lösung Erhalten personalisierter Erlebnisse mit Machine Learning aktualisiert – eine AWS-Lösungsimplementierung, die End-to-End-Automatisierung und -Planung für Ihre Amazon-Personalize-Ressourcen bietet. Diese Lösung hält Ihre Artikel- und Benutzerdaten auf dem neuesten Stand und verwaltet das Retraining Ihrer Modelle, um sicherzustellen, dass die Empfehlungen immer auf dem neuesten Stand der Benutzer-Aktivitäten sind und ihre Relevanz für Ihre Benutzer beibehalten. Diese Lösung veröffentlicht Modell-Offline-Metriken zur Personalisierung auf Amazon CloudWatch, um im Lauf der Zeit eine richtungweisende Orientierung für die Qualität Ihrer Modelle zu liefern.  
Quelle: aws.amazon.com

Bereiten Sie JSON- und ORC-Daten vor, balancieren und codieren Sie Datensätze und launchen Sie Datenverarbeitungsaufträge mit einem Klick mit Amazon SageMaker Data Wrangler

Amazon SageMaker Data Wrangler reduziert den Zeitaufwand für die Zusammenführung und Vorbereitung von Daten für Machine Learning (ML) von Wochen auf Minuten. Mit SageMaker Data Wrangler können Sie den Prozess der Datenvorbereitung und des Feature-Engineerings vereinfachen, und jeden Schritt des Datenvorbereitungs-Workflows, einschließlich der Datenauswahl, -Bereinigung, -Erkundung und -Visualisierung, über eine einzige visuelle Oberfläche abschließen. Mit dem Datenauswahl-Tool von SageMaker Data Wrangler können Sie schnell Daten aus mehreren Datenquellen wie Amazon S3, Amazon Athena, Amazon Redshift, AWS Lake Formation, Amazon SageMaker Feature Store und SnowFlake auswählen.
Quelle: aws.amazon.com