AWS Lambda now publishes logs for Lambda Managed Instances capacity providers

AWS Lambda now publishes logs for Lambda Managed Instances (LMI) capacity providers to Amazon CloudWatch Logs, giving you visibility into scaling activity and instance lifecycle operations. LMI enables you to run Lambda functions on Amazon EC2 instances while maintaining serverless operational simplicity. Capacity providers are resources that let you define compute resources that Lambda provisions on your behalf. With capacity provider logs, you can monitor, troubleshoot, and optimize these managed EC2 instances, helping you quickly diagnose provisioning issues and understand scaling behavior.
Customers use LMI to operate high-volume, predictable workloads with specialized compute configurations and achieve cost efficiency through EC2 pricing options like Savings Plans and Reserved Instances. With this launch, Lambda automatically generates logs for compute resources managed by capacity providers and delivers them to CloudWatch Logs. Lambda publishes structured JSON logs capturing instance lifecycle events like launches, terminations, and health checks. This structured format lets you identify failed operations and provisioning errors through CloudWatch Logs filtering, helping you resolve issues quickly and shorten debugging cycles.
The capacity provider logs are available in all AWS Commercial Regions where LMI is available. The logs are enabled by default for all capacity providers. You can view your capacity provider logs by visiting the Lambda console’s capacity provider page. You can use the Lambda API, Lambda console, AWS CLI, AWS SAM, or AWS CloudFormation to change capacity provider log configuration. Standard Amazon CloudWatch Logs charges apply. To learn more, visit the AWS Lambda Managed Instances product page and documentation. 
Quelle: aws.amazon.com

Amazon Kinesis Data Streams now supports scaling down ingest capacity with warm throughput

Amazon Kinesis Data Streams is a serverless streaming data service that makes it easy to capture, process, and store data streams at any scale. On-demand streams automatically increase ingest capacity in response to rising data ingest usage. With On-demand Advantage mode, you can proactively manage stream capacity using warm throughput to prepare streams for sudden changes in data traffic. We are extending warm throughput with the ability to also scale down ingest capacity, giving you full control to scale your stream’s write throughput up or down.
To scale down, simply set a lower warm throughput value on your on-demand stream. The stream adjusts to the requested capacity or the amount needed to support peak data ingest usage in the last hour, whichever is higher. This ensures your stream always retains sufficient capacity for current traffic while releasing excess capacity you no longer need. As a result, you get optimal stream-processing performance and cost efficiency. 
Warm throughput scale-down is available at no additional cost for all on-demand streams with On-demand Advantage mode enabled.  For more information about On-demand Advantage, see Choose the right mode to stream in in the Amazon Kinesis Data Streams Developer Guide. To get started with the feature, see Update a stream. For pricing details, see Amazon Kinesis Data Streams pricing.
The feature is available in all AWS Regions where Amazon Kinesis Data Streams On-demand Advantage is supported. 
 
Quelle: aws.amazon.com

Amazon EC2 Dedicated Hosts now support host resource groups without self-managed licenses

Starting today, customers can create Host Resource Groups (HRGs) for EC2 Dedicated Hosts without the previously required step of creating Self-Managed Licenses (SMLs) and associating AMIs through AWS License Manager.
This flexibility is particularly valuable for EC2 Mac Instance customers and for customers who need Dedicated Hosts for hardware-level isolation rather than Bring Your Own License (BYOL). Customers with BYOL workloads can continue to create HRGs with SMLs to ensure that only instances from associated AMIs can be launched on the host and track host-level license consumption.
To create an HRG without SML, uncheck the “Restrict to AMIs associated with self-managed license” option when creating a Host Resource Group in the EC2 Console, or set instance-launch-option to license-configuration-required via the AWS CLI.
This feature is available in all AWS Regions where Host Resource Groups are supported. To learn more, visit the Host Resource Group User Guide
Quelle: aws.amazon.com

Amazon MWAA now supports Apache Airflow version 2.11.2

Amazon Managed Workflows for Apache Airflow (MWAA) now supports Apache Airflow version 2.11.2. Amazon MWAA is a managed service that runs Apache Airflow at scale without the operational overhead of managing the underlying infrastructure. Apache Airflow 2.11.2 is a maintenance release that includes security improvements, bug fixes, and dependency upgrades. This release upgrades core dependencies with security patches and stability improvements to the Airflow webserver and task execution layers. It also includes fixes to task lifecycle management for queued tasks, enhanced secrets masking in logs, UI corrections in the Task Instances list view, and provider package updates for S3 and CloudWatch log delivery. You can create a new Apache Airflow 2.11.2 environment on Amazon MWAA or upgrade your existing environments with a few clicks in the AWS Management Console in all currently available Amazon MWAA regions. To learn more, visit the Amazon MWAA documentation, review the Apache Airflow 2.11.2 release notes, and explore the list of available Airflow versions on MWAA.
Quelle: aws.amazon.com

Amazon Connect now supports audio optimization for Azure Virtual Desktop and Windows 365 Cloud PC

Agents using Microsoft Azure Virtual Desktop (AVD) or Windows 365 Cloud PC can now take calls directly from their virtual desktop session with audio optimization enabled. To get started, IT administrators need to complete a one-time setup for their virtual desktop environment. Once configured, media is redirected from the virtual desktop to the agent’s local device, improving audio quality. Agents simply log into their Azure Virtual Desktop or Windows 365 Cloud PC session and start accepting calls using the Amazon Connect Customer agent workspace or a custom agent interface built with the Amazon Connect Customer open-source JavaScript libraries. This support is in addition to existing audio optimization for Amazon WorkSpaces, Citrix cloud desktops, and Omnissa cloud desktops. This feature is available in all AWS Regions where Amazon Connect Customer is offered, except AWS GovCloud (US-West). To learn more, see the Amazon Connect Customer Administrator Guide.
Quelle: aws.amazon.com

AT&T and Microsoft scale trillion-token workloads with Microsoft Foundry and AMD

First-of-its-kind telecom AI deployment

Telecommunications organizations are increasingly looking to AI to help teams navigate highly specialized domains, but generic models often lack the industry-specific knowledge needed to understand telecom networks, standards, and operations. To address that gap, AT&T created their Open Telco (OTel) models, the next generation of telecom-focused AI designed to bring deeper telecommunications expertise into AI systems. Building OTel2.0 required more than training a large language model, it reflected a broader issue many organizations face: how to build domain-specific AI systems at scale while balancing cost, performance, and operational complexity. Cost management quickly became a key consideration. To continue advancing telecom-focused AI, AT&T needed a platform capable of supporting OTel2.0 development at an entirely new scale.

Where teams previously had to own and manage deployments, infrastructure, and the associated operational overhead, Foundry Managed Compute provided a more streamlined way to access dedicated graphics processing unit (GPU) capacity. This transformation requires more than powerful models; it requires the ability to scale without compromising cost, flexibility, or performance.

Learn how OTel2.0 scales telecom AI

Using Microsoft Foundry Managed Compute, AT&T was able to experiment across multiple open models, optimize workloads across different GPU architectures, and process massive volumes of telecom data all within a unified platform. The result was an AI development environment capable of supporting trillions of tokens while giving teams the flexibility to iterate, optimize, and innovate faster.

Model choice meets infrastructure flexibility

Building OTel2.0 required flexibility across both models and infrastructure. Rather than standardizing on a single model, AT&T adopted a multi open-model strategy. Open models were central to AT&T’s approach because they provided the flexibility to work with approved telecom data, tailor the workflow for domain-specific model development, and support large-scale experimentation with greater control over cost and deployment strategy. Through Microsoft Foundry, the team deployed several models from the Hugging Face collection, including Phi-4, OSS-120B, and Gemma-4, to support different stages of development, from synthetic data generation and data preparation to reasoning-intensive workloads and broader model development efforts. Phi-4 played a significant role in this process, processing more than 700 billion tokens a month as part of the broader data preparation and training workflow for OTel2.0.

Every company in the world needs to build its own AI, and that is only possible with open models and open source. AT&T is championing this vision, building on open models like Phi-4 and Gemma, and giving OTel back to the community as a telecom AI foundation others can build upon. Microsoft Foundry makes this practical at scale, bringing the latest open models from the Hugging Face collection together with AMD and NVIDIA GPUs in one place, so teams can pick the right model and the right hardware, then deploy in hours instead of weeks.
—Jeff Boudier, Vice President of Product, Hugging Face

Developing OTel2.0 also required infrastructure capable of operating at telecom scale. AT&T used approximately 530 GPUs through Microsoft Foundry Managed Compute spanning multiple GPU architectures including 430 AMD Instinct™ MI300X GPUs. This heterogenous approach gave AT&T more flexibility in how models were deployed and optimized as requirements evolved.

ModelExample workloadPhi-4Around 700B tokens a month for data preparation and synthetic data generationOSS 120BHigher-reasoning workloads Gemma 4OTel2.0 development workflowsTable 1: Explains what open source models were used and how

This flexibility illustrates a broader trend across AI development. Organizations increasingly need platforms that allow them to choose the right model for the job, optimize for cost and performance, and scale workloads without rebuilding operational environments. Microsoft Foundry brings model choice, infrastructure flexibility, governance, and operational scale together in a unified platform that supports those requirements.

Start building with Microsoft Foundry

Beyond flexibility and cost, deployment speed is a critical factor for many AI initiatives. As workloads expand and new models are evaluated, the ability to access GPU capacity quickly enables teams to move from experimentation to execution faster without lengthy provisioning cycles. With Foundry Managed Compute, AT&T could deploy and scale models in days rather than waiting weeks for infrastructure to become available, helping accelerate development timelines and maintain momentum across OTel2.0 development.

Optimizing cost without limiting innovation

As AI workloads grow, economics become as important as model performance. For AT&T, one of the primary objectives was to lower AI model consumption costs while continuing to drive meaningful business value through AI-powered innovation. By using open models on Microsoft Foundry Managed Compute, AT&T was able to support large-scale data preparation and model development using a different economic model built around dedicated GPU infrastructure and open-model flexibility.

The impact became clear at scale. In support of OTel2.0, AT&T processed approximately 1T tokens, consisting of raw documents from GSMA supplemented by synthetic data generated. Generating the data using open-source models like Phi-4, served by Microsoft’s Foundry Managed Compute, saved tens of millions of dollars versus using frontier models. This allowed teams to invest in larger-scale experimentation and development while maintaining a focus on business value and operational efficiency.

MetricValueOTel 1.0 DownloadsOver 25MGPUs Used Through Foundry Managed ComputeAbout 530Tokens Processed for OTel2.0About 1TTokens Trained for OTel2.0About 400 BModels used to train OTelPhi-4, OSS 120B, Gemma 4Table 2: Quick facts about the OTel model family and metrics around what was used to build OTel2.0 

When you are processing hundreds of billions of tokens, infrastructure becomes part of the problem you solve. Foundry Managed Compute gave us access to GPU capacity at scale so our teams could focus on advancing OTel2.0 instead of managing infrastructure.
—Mark Austin, Vice President, Data Science and AI at AT&T

At this scale, infrastructure is no longer simply a deployment consideration. It becomes a strategic component of AI development.

Accelerating the next wave of production-scale AI

OTel 2.0 demonstrates how organizations can combine open models, scalable infrastructure, and domain expertise to build production-ready AI systems. By matching different models to different workloads and optimizing infrastructure for cost and performance, AT&T was able to process trillions of tokens while maintaining operational efficiency. 

As organizations move from AI experimentation to production deployment, they increasingly need the flexibility to choose the right models, optimize infrastructure, and scale efficiently. Microsoft Foundry and Foundry Managed Compute help support that transition by bringing those capabilities together in a unified platform.

Learn more

Read Scott Guthrie’s blog about Azure AI and HPC infrastructure.

Learn more about OTel2.0.

Explore session topics from AMD’s Advancing AI:

From GPUs to CPUs: Optimizing Every AI Workload with Azure and AMD

What’s Next for AI Infrastructure in the Cloud?

Powering the Future of AI on Azure

Discover how Microsoft and AMD are expanding Azure AI and HPC infrastructure.

Read more

The post AT&T and Microsoft scale trillion-token workloads with Microsoft Foundry and AMD appeared first on Microsoft Azure Blog.
Quelle: Azure