Leave manual cluster resizing behind with Cloud Dataproc’s autoscaling

Building real-time, interactive data products with open source data and analytics processing technology is not a trivial task. It involves constantly balancing cluster costs with service-level agreements (SLAs). Whether you are using Apache Hadoop and Spark to build a customer-facing web application or a real-time interactive dashboard for your product team, it’s extremely difficult to handle heavy spikes in traffic from a data and analytics perspective.We’re pleased to announce Cloud Dataproc’s new autoscaling capabilities, now generally available, that can remove the need for complex capacity planning that always results in either missed SLAs or resources sitting idle.How can autoscaling help your team?These new capabilities can help a range of teams, whether data engineers building complex ETL pipelines, data analysts running ad hoc SQL queries, or data scientists training a new model. Cloud Dataproc’s autoscaling capabilities allow cluster admins to build ephemeral or long-standing clusters in 90 seconds and apply an autoscaling policy to the cluster to minimize costs and maximize the user experience without manual intervention. Whether you’re part of the team at a technology company building a SaaS application, a telecommunications company analyzing network traffic, or a retailer monitoring clickstream data during the holidays, you no longer have to worry about right-sizing clusters. Here’s a look at some common use cases:Core Cloud Dataproc autoscaling capabilities include:Right-sizing your cluster: Estimating the “right” number of cluster workers (nodes) for a workload is difficult, and a single cluster size for an entire pipeline is often not ideal. Don’t worry about manually right-sizing your cluster with autoscaling. One autoscaling policy, multiple clusters: An autoscaling policy is a reusable configuration that describes how clusters using the autoscaling policy should scale. It defines scaling boundaries, frequency, and aggressiveness to provide fine-grained control over cluster resources throughout the cluster lifetime.Budget optimization: Scale in and scale out clusters while setting limits in the autoscaling policy to make sure you don’t exceed budget. YARN integration:Autoscaling policies integrate with YARN automatically to trigger VM scaling when needed, so you have one central resource management system for all of your Cloud Dataproc jobs.Monitor autoscaling jobs: Integrate with Stackdriver Monitoring to view the metrics from the autoscaling clusters, view the number of Node Managers in your cluster, and understand why autoscaling did or did not scale your cluster. Use Stackdriver Logging to view autoscaler decisions.Multi-region support: Deploy autoscaling clusters in any region where Cloud Dataproc clusters are running. Check out our documentation to access everything you need to get started with Cloud Dataproc autoscaling. Autoscaling is supported through the v1 API on cluster image versions 1.0.99+, 1.1.90+, 1.2.22+, 1.3.0+, and 1.4.0+.
Quelle: Google Cloud Platform

Einführung von AWS Data Exchange

AWS Data Exchange ist ein neuer Service, der es Millionen AWS-Kunden erleichtert, Daten von Drittanbietern in der Cloud sicher zu finden, zu abonnieren und zu benutzen. Qualifizierte Datenanbieter schließen die Folgenden ein: Reuters, Foursquare, TransUnion, Change Healthcare, Virtusa, Pitney Bowes, TP ICAP, Vortexa, IMDb, Epsilon, Enigma, TruFactor, ADP, Dun & Bradstreet, Compagnie Financière Tradition, Verisk, Crux Informatics, TSX Inc., Acxiom, Rearc und viele mehr.  
Quelle: aws.amazon.com

Bring Your Own IP für Amazon Virtual Private Cloud ist jetzt in fünf zusätzlichen Regionen verfügbar

Bring Your Own IP (BYOIP) steht ab heute in den folgenden AWS-Regionen zur Verfügung: Asien-Pazifik (Mumbai), Asien-Pazifik (Sydney), Asien-Pazifik (Tokio), Asien-Pazifik (Singapur), Südamerika (São Paulo) und zusätzlich in EU (Dublin), EU (London), EU (Frankfurt), Kanada (Zentral), USA Ost (Nord-Virginia), USA Ost (Ohio), und USA West (Oregon). BYOIP wird ab heute ebenfalls die Einbindung von IP-Adressen, die im Asia Pacific Network Information Center (APNIC) registriert sind, zusätzlich zu denen im American Registry for Internet Numbers (ARIN) und Réseaux IP Européens Network Coordination Centre (RIPE) unterstützen.
Quelle: aws.amazon.com

Amazon ElastiCache unterstützt jetzt T3-Standard Cache-Knoten

Sie können jetzt die neue Generation der burstfähigen T3-Standard Cache-Knoten für allgemeine Zwecke in Amazon ElastiCache starten. Amazon EC2s T3-Standard-Instances bieten eine solide CPU-Leistung und die Möglichkeit, diese Leistung jederzeit, und so lange bis das angesammelte Guthaben aufgebraucht ist, zu bursten. Sie bieten Generationsvorsprünge in CPU-Leistung, die eine höhere gesamte Grundschwellenleistung über T2 Cache-Knoten ermöglichen. 
Quelle: aws.amazon.com

Introducing Batch on GKE—modernizing HPC with Kubernetes in the cloud

One of our most important goals at Google Cloud is to make cloud computing easier, so that you can focus on answering questions that matter the most to you, your business and your users—not on managing infrastructure. Today, we are excited to announce the preview of Batch on Google Kubernetes Engine (GKE), a cloud-native solution for running batch workloads at scale in an optimized manner. Batch on GKE brings the functionality and familiarity of a traditional batch job scheduler into a cloud-first world. It frees your applications from the limitations of fixed-sized compute clusters by dynamically allocating resources to meet the needs of your application.Google Cloud is the home of Kubernetes—originally developed here and released as open source in 2014. GKE is a managed, production-ready environment for deploying containerized applications that relieves your teams from the operational toil of Kubernetes cluster management, allowing you to focus on your business needs. We heard you wanted to bring the benefits of GKE to batch workloads such as media rendering, genomics sequencing, silicon design verification and financial portfolio risk analysis, so we built Batch on GKE to bring what you love about GKE to these real-world batch workloads.The preview release of Batch on GKE comes with the following capabilities:Autoscaling and just-in-time provisioning to ensure you pay for just what you need Rightsizing of virtual machines to tailor fit CPU and memory for the job at hand Smart reuse of virtual machines and smart packing of jobs to reduce waste and the time jobs spend waiting in a queueResource budgets to allocate the maximum spend per teamGraphics Processing Unit (GPU) supportJob submission tool that high performance computing practitioners will find familiarBatch on GKE’s easy-to-use, familiar interface enables you to deliver business results for your batch computing use cases. The following diagram shows how users can leverage Batch on GKE and other Google Cloud services to build a genomics processing and analysis system.Click to enlargeMeeting you where you are with partnersBatch on GKE is a great solution for modernizing your batch workloads. We’re also committed to helping you migrate your existing systems as-is to Google Cloud or augment your on-premise setup by connecting to Google Cloud. We partner with SchedMD, Altair and Univa to integrate their market-leading schedulers with our platform and meet you where you are.  Engineering simulation made easyWe’re also working closely with Rescale to enable their full-service HPC platform to run engineering simulations on Google Cloud and leverage on-demand GCP clusters and virtually unlimited cores. With this integration, you can run and manage simulations from a vast application library, including ANSYS Fluent, LS-DYNA, Star-CCM+ and more. Click here for a list of all the software that Rescale supports on Google Cloud. If you’re attending SuperComputing 2019, be sure to visit the Google Cloud booth (#1363). You can find more information about our SC19 presence here. You can also learn more about how Google Cloud’s flexible infrastructure can accelerate your HPC workloads.Special thanks to Senanu Aggor, Product Marketing, and Annie Ma-Weaver, Partner Manager, for making this blog post possible.
Quelle: Google Cloud Platform