Nintendo erwirkt Versäumnisurteil: Switch-Reparatur führt zu Schadensersatzklage
Über eine angeforderte Hardware-Reparatur hat Nintendo ein auffälliges Reddit-Profil mit einer Wohnadresse verknüpft. (Nintendo, Urheberrecht)
Quelle: Golem
Über eine angeforderte Hardware-Reparatur hat Nintendo ein auffälliges Reddit-Profil mit einer Wohnadresse verknüpft. (Nintendo, Urheberrecht)
Quelle: Golem
Elf Merchandising-Artikel bekommen GTA-6-Fans in der neuen Goodtime-State-Vice-City-Collection von Rockstar. Das Spiel ist nicht dabei. (GTA 6, Rockstar)
Quelle: Golem
Die Apple Airpods Pro 3 In-Ear-Kopfhörer fallen erstmals seit Längerem bei Amazon wieder unter die 200‑Euro-Marke. (Airpods, Apple)
Quelle: Golem
Anwender sollen Spiele für Meta-Plattformen mit KI direkt am Smartphone und im Browser prompten können. (KI, Softwareentwicklung)
Quelle: Golem
Latin America’s small and medium-sized businesses are the heartbeat of the region’s economy — accounting for more than 60% of total employment in the region, according to United Nations estimates. And just like their enterprise peers, everywhere you look, ambitious teams are moving fast to embrace AI.
Many have already transitioned from experimenting with generative tools and agentic workflows to using them every day to work smarter, save time, and deliver exceptional customer experiences. These growing businesses are particularly focused on maximizing the benefit they get from their investments in AI, whether that’s using a fast, low-cost model to summarize daily emails or deploying an advanced model for complex data analysis, teams can match the right AI capability to their exact task and budget.
It’s this range of options, and a familiarity with the broader suite of Google business, media, and advertising tools that has led many SMBs to choose Google Cloud, and Gemini Enterprise in particular, as their AI platform of choice. By doing so, they’re able to build custom AI agents, streamline daily tasks and paperwork, and offer customers instant support with the speed and reach needed to compete on a global scale.
With our unique front row seat, we’ve seen the benefit SMBs are getting from leveraging Gemini Enterprise, not only for generative AI, but as a catalyst for adopting other essential cloud tools like Google Kubernetes Engine and BigQuery for complete end-to-end modernization. The number of Latin American-based small and medium businesses using Google Cloud AI tools has grown 8x year-over-year and the number of Brazil based small and medium businesses using Google Cloud AI tools has grown year-over-year. This rapid adoption spans our Gemini models, Gemini Enterprise, and core Cloud infrastructure, and are helping businesses to:
Roll out better customer support systems to help escalate and resolve customer support calls more quickly.
Automate repetitive actions in areas like payroll and accounting.
Help more employees understand and leverage data at work — even those not trained as data analysts.
Rapidly create and implement new designs for marketing collateral.
Help more people build their own AI agents to help them in their everyday jobs.
As we head into today’s Google Cloud Summit in Brazil, we were proud to showcase nearly 20 of our newest Latin American SMB customers using Google AI to reduce busywork, serve their customers faster, and grow their businesses.
Announcing new Latin American customers putting Google AI to work
AdGoat, an Argentina-based adtech company processing more than 10 billion annual ad requests across more than 100 global websites. It uses Cloud Run, the Gemini API, and Gemini Enterprise to automate content analysis, ad bidding, and audience targeting to help e-commerce brands drive higher campaign returns.
Angelus, a Brazilian dental and healthcare manufacturing company, uses Gemini Enterprise to streamline project management across its research and development department. This enables its teams to automatically pull technical project data into pre-approved templates aligned with the company’s brand identity and regulatory requirements.
BunkerDB, a marketing science company operating across Latin America, uses Gemini Enterprise, Cloud Run, and Cloud Storage to power an AI platform that organizes marketing assets, checks brand compliance, generates or adapts multimodal content, and predicts ad performance before launch. All of this helps it reduce creative turnaround times from weeks to hours and cut cost per lead by up to 25%.
Caffeine Army, a Brazilian wellness and high-performance company that connects people with solutions in nutrition, sports, and well-being, deployed BigQuery and Gemini Enterprise on Google Cloud to unify customer purchase insights, enabling faster creative campaign turnarounds and boosting team productivity across the organization.
Convert, a Brazilian marketing and analytics provider, uses Looker, BigQuery, and Cloud Run to power five specialized AI agents that answer complex business questions in natural language, speeding up report deliveries by 65% and reducing operational costs by 32%.
Growth Digital, a Google Ad sales rep operating across 13 Latin American countries, used BigQuery and Gemini Enterprise to build over 113 AI agents, enabling teams to build proposals 5x faster, cut campaign reporting time by 80%, and reduce financial error rates to under 0.01%.
GrupoTusMaquinas.com, an equipment management platform based in Chile, deployed Google Cloud AI tools and Gemini models to create digital tracking profiles for trucks and machinery, allowing businesses to query fleet status in plain language and manage vehicles regardless of brand or location.
HealthAtom, a healthcare technology company, uses the Gemini API, Firestore, and Cloud Functions to power AI assistants across its clinical platforms, automating appointment scheduling and medical record reviews while supporting 80 million annual patient interactions.
KLog.co, a Chilean logistics technology company digitizing freight forwarding across Latin America, uses Gemini Enterprise, BigQuery, and Google Workspace to automate cargo tracking and shipping paperwork, cutting manual data entry errors by over 90% and increasing document processing capacity tenfold.
NEEOH, a leading Brazilian out-of-home advertising communication platform, uses Gemini Enterprise to standardize secure AI usage across its organization, enabling teams to generate campaign copy and build pitch proposals faster while keeping corporate client data secure.
Luxia Agro, an Argentinian foreign trade supplier of crop protection products, uses Gemini 3.5 Flash and Gemini Enterprise to automatically pull key details from complicated shipping emails and update their central business systems. This allows it to automate 80% of foreign trade operations and cut manual processing errors in half.
Macal, a Chilean auction company, uses the Gemini Enterprise, Cloud Run, BigQuery, and Security Command Center to automatically verify property records and modernize its technology systems, cutting software development times from weeks to days and lowering infrastructure costs by up to 30%.
Ninecon, a Brazilian tech consulting firm, deployed Gemini Enterprise to integrate AI directly into employee workflows, allowing managers to track usage patterns and optimize project turnaround times with real-time insights.
Nova Gestões, a customer service and operations provider in Brazil, uses Cloud Speech-to-Text and Gemini Enterprise to translate and analyze 100% of customer calls in real time, reducing post-call manual data entry and boosting team productivity by 30%.
Romi, a Brazilian industrial machinery manufacturer, uses the Gemini API and Gemini Enterprise to power an interactive chat assistant directly on CNC machine HMI (human machine iInterface), giving factory operators instant answers grounded in official manuals and generating QR codes for step-by-step instructional videos.
Supermercados El Dorado, a leading supermarket chain in Uruguay, leverages Google Compute Engine and Gemini Enterprise to modernize legacy testing infrastructure and connect custom AI agents within their daily workflows, boosting team productivity across departments.
Tryvia, a Brazilian IT and business solutions provider, uses Google Cloud, Looker, and Gemini Enterprise to move off legacy physical servers, giving teams real-time reporting dashboards and AI tools that speed up software development and daily tasks.
Via Cristais, a major highway operator in Brazil, leverages Google Contact Center as a Service to speed up emergency routing for highway accidents, reducing caller wait times, improving driver satisfaction, and mitigating the impact of call center staff turnover.
WeSpeak, an AI conversational platform for the hospitality industry in Latin America, uses Cloud Run, Gemini Pro, and Gemini Flash to automate end-to-end guest interactions across messaging channels like WhatsApp and Instagram. This has helped it achieve an 85% resolution rate and a 2x increase in overall sales volume for hotel clients.
Helping your team build AI skills
To help growing teams get the absolute most out of AI, we’ve created easy, no-cost learning programs that anyone can use:
Programs for small and medium businesses: Explore beginner-friendly training paths or join specialized programs to learn how to build custom AI assistants for your day-to-day work.
Google skills for organizations: Access thousands of free, on-demand AI courses and hands-on practice labs designed by experts at Google Cloud and Google DeepMind.
Get certified: Help your staff gain industry-recognized AI certificates through guided courses, expert mentoring, and skill badges.
By offering easy-to-use tools and free training — from everyday office apps in Workspace to advanced AI on Google Cloud — Google is here to help Latin American businesses thrive today and in the future.
Quelle: Google Cloud Platform
AI models are increasingly taking on work that extends far beyond a single prompt: building a feature across a codebase, investigating a complex issue, synthesizing hundreds of pages of information, or working through a multi-step business process.
As that work gets longer, raw intelligence is only part of what matters. The model also needs to stay focused, make good decisions along the way, communicate what it is doing, and produce work that people can quickly review and use. Today, Claude Opus 5.5 is available in Microsoft Foundry, bringing Anthropic’s most capable Opus model to developers and enterprises building AI applications and agents.
Claude Opus 5.5 is designed for everyday complex work. It advances Opus 5 across agentic coding, knowledge work, and long-running tasks while making it easier for people to understand what the model did, what it found, and what it needs next. Claude Opus 5.5 also does more with fewer tokens. Lower per-token prices and much cheaper cache reads stack on top of the efficiency gains.
Built for work that takes time
Writing a function is one thing. Building a feature that touches multiple services, tracing a production issue across a large repository, or carrying a task from investigation through implementation and validation is something else entirely. Claude Opus 5.5 is designed for these longer-running workflows.
For software development, it can work through long-running coding tasks such as building features across a codebase, debugging, refactoring, and reviewing code. It finds the root cause before changing anything, checks its work as it goes, and explains its changes in plain language, so engineers can review and trust them quickly.
That combination becomes particularly valuable when developers use models through agentic coding environments, where a session may involve dozens of steps and run for an extended period of time. The same applies beyond software development.
For knowledge workers, Claude Opus 5.5 can bring together information from multiple sources, work through long documents and spreadsheets, perform analysis, and help create artifacts such as memos, reports, and presentations. Compared with Opus 5, it produces outputs that require less editing before they are ready to share.
An AI model that communicates more like a teammate
As agents take on more autonomous work, another challenge emerges: keeping the human in the loop without overwhelming them.
An agent that performs 50 steps should not require someone to inspect 50 steps to understand whether the work was successful. Claude Opus 5.5 introduces improvements to agentic communication designed to make long-running work easier to follow. As it works, the model can surface the information that matters most:
What it did
What it found
What decisions it made
Where it needs input from the user
What happened at the end of a long-running task
The goal is simple: spend less time decoding what the model did and more time using the result. This matters particularly for enterprise agents, where users need to understand not only the final answer but also when an agent needs clarification, encounters a constraint, or reaches a decision point that requires human judgment.
Adpative thinking
Claude Opus 5.5 uses adaptive thinking, automatically determining how much reasoning a task requires. Rather than turning thinking on or off or manually specifying a thinking-token budget, developers use effort to influence how much work the model should put into a request. This allows the model to adapt its reasoning to the task at hand—from relatively straightforward requests to complex problems that require deeper analysis.
For developers building agents, this can reduce the amount of application logic needed to decide when and how a model should reason.
Designed for long-running agent architectures
Long-running agents create challenges beyond model intelligence. Conversations can exceed context limits. Tools available to an agent can change. Applications may need to compact earlier context while preserving the model’s understanding of the work already completed.
Alongside Claude Opus 5.5, Anthropic is introducing beta API capabilities designed for these scenarios, including asynchronous compaction, keep-tail compaction, and changing tools during a conversation while preserving thinking and prompt caching. These capabilities can help agent developers maintain continuity across longer tasks without rebuilding the state of the application every time context or available tools change.
Combined with Microsoft Foundry, developers can use Claude Opus 5.5 as part of broader agent systems that connect models with enterprise data, tools, evaluation, and operational workflows.
Expanded safeguards for more capable models
As model capabilities increase, Anthropic is also expanding the safeguards applied to Claude Opus 5.5.
Claude Opus 5.5 is the first Opus model to use safety classifiers like those introduced with Claude Fable 5.1 in areas including cybersecurity, biology, AI development, and distillation.
For common developer, educational, and knowledge-work scenarios, customers can continue using the model for tasks such as identifying software vulnerabilities or learning about biological concepts. Certain requests that Anthropic identifies as higher-risk or dual-use may be handled by another Claude model with the appropriate safeguards.
This reflects an increasingly important part of deploying more capable models: advancing what models can do while applying safeguards appropriate to the capabilities they introduce.
Pricing
ModelDeployment OffersInput/M TokensOutput/M TokensAvailabilityClaude Opus 5.5Global Standard, US DataZone$4Cache Hit – $0.20Cache Write – $5Cache Write (1 hr) – $8$20GA, Hosted on AzureClaude Opus 5.5 (Long Context)Global Standard, US DataZone$4Cache Hit – $0.20Cache Write – $5Cache Write (1 hr) – $8$20GA, Hosted on Azure
Build with Claude Opus 5.5 in Microsoft Foundry
Choosing a model is only the beginning of putting AI into production. Microsoft Foundry gives developers a unified place to discover models, build and evaluate AI applications and agents, connect them with enterprise data and tools, and operate those systems in production. As models become capable of taking on more complete units of work, the question is shifting from Can the model answer this prompt? to Can I trust it to carry the work forward? Claude Opus 5.5 represents another step in that direction: stronger performance on complex work, more adaptive reasoning, and clearer communication between people and the AI systems working alongside them.
Claude Opus 5.5 is available today in Microsoft Foundry.
Updated Sep 22, 2026
Version 2.0
The post Claude Opus 5.5 comes to Microsoft Foundry for long-running coding and knowledge work appeared first on Microsoft Azure Blog.
Quelle: Azure
Today, we are expanding our GPT-6 series by welcoming GPT-6 Sol and GPT-6 Luna to our generally available lineup in Microsoft Foundry. Building on the exceptional customer momentum of GPT-5.6 Sol and GPT-6 Astra, this launch continues our work to deliver transformative capabilities in Microsoft Foundry that produce less noise and are more capable at completing full tasks with agents.
Explore GPT models in Foundry today
Astra brings advanced reasoning, software engineering and computer use to demanding work that requires both judgment and action. Azure customers report a step-change in capabilities, and strong cost-to-performance with the model using fewer, higher-value tokens to drive agents.
Completing the lineup, GPT-6 Sol is excellent for general-purpose use, while Luna brings efficient intelligence to high-volume data and preparatory work.
Put the right intelligence behind every agent
The right model for a job should be determined through evaluations: an agent handling a complex business decision and one routing routine requests have different needs. Microsoft recommends customers start with GPT-6 Astra for demanding work. For higher-volume workloads, GPT-6 Sol and Luna carry that progress forward, giving you a complementary choice built for production and scale.
GPT-6 Sol for production AI agents and complex workflows
GPT-6 Sol, and its proven predecessor—GPT-5.6 Sol—offer slightly more cost-effective intelligence with frontier efficiency. They support enterprise agents, coding and complex knowledge work, including reasoning across multiple steps, long-context analysis, and workflows that use tools. For teams evaluating their next production workload or migrating off a legacy model, Sol is a strong starting point.
GPT-6 Luna for efficient, high-volume AI workloads
GPT-6 Luna is Sol’s smaller, faster sibling, built for high-volume work. Use it for extraction, summarization, request routing, and routine customer interactions. Reserve deeper reasoning for the steps that need it, rather than applying the same model to every task.
As the GPT-6 lineup expands, the opportunity is not simply to choose a newer model, but to improve what your agents can accomplish while saving money. Customers should look beyond pricing per token and seek to understand cost per task, which is a better measure for understanding the ROI of AI.
The accompanying chart illustrates why enterprise customers on Microsoft Foundry are switching to GPT-5.6 Sol and the latest GPT-6 offerings.
Foundry brings evaluation and monitoring together so teams can make those decisions with evidence. The real measure of that progress is what customers can do in production, which is why Foundry has always encouraged model choice and an open, interoperable stack.
The Foundry advantage, in customers’ words
Access to frontier models is only the starting point. Foundry pairs GPT-6 intelligence with the breadth of deployment options enterprise production demands. Today, Standard deployment is available for Astra, Sol and Luna across all 28 Global regions, and US and EU Data Zones; Provisioned Throughput for Astra and Sol across Global regions and US and EU Data Zones; and Priority Processing for Sol across Global regions and US Data Zones. The breadth and performance of Azure is why OpenAI continues to launch first on Azure, and why sophisticated customers like Manus choose Foundry.
Azure OpenAI models provide a core layer of intelligence powering Manus. Through Azure, we reliably integrate advanced models into our agentic workflows, enabling Manus to understand user intent, plan tasks, and execute complex work. Responsive Microsoft technical support and rapid access to new model capabilities help us iterate quickly and deliver a leading, reliable AI experience for our users.
—Tao Zhang, Co-Founder & Product Partner, Manus
For customers getting started with AI on Azure: choose Global for flexible, pay-per-token capacity, or supported Data Zone deployments for processing-location requirements. Priority Processing is a priority lane for responsive, pay-as-you-go experiences, with Provisioned Throughput providing reserved capacity and superior latency for critical production demand. Match the serving option to the workload, from interactive agents to high-throughput business processes.
That is the Foundry advantage: not just frontier intelligence, but the platform to put it to work. Teams can match each workload to the right model, deployment option, and controls, balancing capability, responsiveness, and cost as adoption grows. By bringing these choices together on Azure, Foundry helps customers focus on delivering business value, with the operational foundation to move from a promising agent to production at scale.
Our customers work in domains where getting an answer isn’t enough, it has to be the right answer, and it has to hold up to scrutiny. The latest Azure OpenAI frontier models reason through a problem in steps we can follow, which is what makes it viable for the research and compliance workflows our professionals depend on. Building on Microsoft Foundry lets us take those agentic workflows into production on infrastructure and services that already meet our governance, data residency, and security obligations.
—Brian Diffin, CTO of Wolters Kluwer Tax & Accounting
GPT-6 pricing and deployment options**
ModelDeploymentContext LengthPricing (USD $/million tokens)InputCached InputCached WritesOutputGPT-6 AstraGlobal StandardShort context$10.00$1.00$12.50$50.00Long context$20.00$2.00$25.00$75.00Data Zone Standard (US)Short context$11.00$1.10$13.75$55.00Long context$22.00$2.20$27.50$82.50Data Zone Standard (EU)Short context$12.00$1.20$15.00$60.00Long context$24.00$2.40$30.00$90.00GPT-6 SolGlobal StandardShort context$2.00$0.20$2.50$10.00Long context$4.00$0.40$5.00$15.00Data Zone Standard (US)Short context$2.20$0.22$2.75$11.00Long context$4.40$0.44$5.50$16.50Data Zone Standard (EU)Short context$2.40$0.24$3.00$12.00Long context$4.80$0.48$6.00$18.00GPT-6 LunaGlobal StandardShort context$0.10$0.01$0.125$0.50Long context$0.20$0.02$0.25$0.75Data Zone Standard (US)Short context$0.11$0.011$0.1375$0.55Long context$0.22$0.022$0.275$0.825Data Zone Standard (EU)Short context$0.12$0.012$0.15$0.60Long context$0.24$0.024$0.30$0.90
**Pricing for both Provisioned Throughput and Priority Processing varies by deployment type. For each offer, U.S. Data Zone is priced at a 10% premium to Global. For current rates and terms, see the Azure OpenAI pricing page.
Build safer AI agents with Microsoft Foundry
GPT-6 models running on Azure have multiple layers of safety and security built directly into the model and around it. At the core, the model itself carries the alignment and safety training built in, while the prompts and outputs around it are protected by content filters and guardrails that govern what the agent can say. Beyond that, tool calls and responses are protected by controls and prompt injection mitigation that govern what the agent can do, and identity and access are protected by enterprise policies that govern what it can reach.
Foundry helps teams continuously strengthen safety layers as risks evolve. It applies guardrails at key checkpoints, including prompts, outputs, tool calls, and tool responses. Identity and access controls govern what agents can do and reach. Microsoft Purview applies enterprise data policies. Evaluation, tracing, and monitoring give teams the evidence to optimize those controls over time, with human checkpoints at every phase.
Move to GPT-6. Build your next generation of agents.
Build your next agentic workloads in Microsoft Foundry. Start with GPT-6 Astra for demanding reasoning, Sol for general production use, and scale high-volume tasks with GPT-6 Luna. For customers of legacy models, we recommend evaluating an upgrade to GPT-5.6 Sol and above.
Your next agent needs more than a powerful model. Foundry brings an open intelligence stack, deployment flexibility, and Azure enterprise controls together so you can build with confidence and scale from your first workload to production.
Start building in Foundry today
Access GPT-6 models, evaluate the right fit for your workload, and scale from experimentation to production.
Begin here
The post GPT-6 Astra, Sol, and Luna: For production agents in Microsoft Foundry appeared first on Microsoft Azure Blog.
Quelle: Azure
For decades, applications have been designed to wait. A user clicks, a request arrives, code runs, a response goes back. And we got very good at this. We learned to forecast traffic, scale on demand, wrap everything in enterprise controls, and run it all with the operational discipline that keeps critical business systems available around the clock.
That model is being turned on its head, and being asked to serve apps that continuously act, and wait for no one.
What changed is not the infrastructure underneath, but the software being written on top of it. The work we used to capture as deterministic code, where every branch was defined in advance and every step was known before the first line ran, is now being rewritten as multi-agent applications that work out the steps at runtime. A developer used to encode the path, and now a developer describes the outcome and lets a set of agents reason their way toward it.
Agents work differently. Given an outcome, an agent reasons through the problem, breaks it into steps, writes code to solve it, runs that code, looks at the result, and goes again. The work happens in a loop that no human is standing inside. That change puts pressure on assumptions our platforms were built on, because an application that responds and an agent that acts need very different things underneath them.
The organizations pulling ahead right now are the ones who understood this early. They are not simply adding AI to what they already have. They are designing for a different kind of software.
What it takes to run an agent you can depend on
Getting an agent working is no longer the hard part. A team can connect a capable model to a few tools, ground it in company data, and have something genuinely useful inside a week. That is real progress, and it is why so many organizations now have a pilot that impressed the room.
The distance between that pilot and an agent the business depends on around the clock is where the real challenge is.
An agent operating continuously on behalf of a company gets held to the same bar as everything else in production. It needs its own identity, with permissions scoped to what it is allowed to see. It needs to be watched while it works, so a team can trace every step it took, evaluate whether the outcome was right, and catch the drift that shows up quietly weeks after launch. It needs guardrails that are enforced at runtime. And it has to clear the same security and compliance standards as the rest of the estate. No one is going to grant an agent an exemption.
None of that is work a team should be doing for itself. That is the job of an agent platform. Microsoft Foundry is where agents are built, grounded in enterprise knowledge, given a first-class identity through Entra Agent ID, and traced and evaluated once they are live. The Foundry Control Plane governs the agent, but it does not dictate where the agent’s work actually runs. That is a separate decision, and a separate layer.
Get started with Microsoft Foundry
Where the work actually runs
The moment an agent stops answering questions and starts completing tasks, it has to execute code. It clones a repository. It installs a package. It runs an analysis against live data and calls into a system of record. That execution needs a runtime environment. In most cases, it simply inherits the one the host application is already running on, because that is the path of least resistance.
Run the agent on shared infrastructure and every workload inherits the blast radius of every other one. Give it broad access so it can be useful and you have handed an autonomous process far more reach than you intended. Lock it down until it is safe and the agent can no longer do the job you built it for. Teams get stuck in this trade-off, and it is where most promising agents usually stall before they reach production.
The answer is not to limit what the agent can do. It is to give it a dedicated environment, with its own identity and guardrails for execution.
Azure Container Apps Sandboxes provides exactly that. Every agent execution gets a fully isolated environment of its own, created in seconds and gone when the work is finished. It runs as an identity you control, can access only the systems you have allowed it to reach, and never stores the credentials it uses to get there. And when a task spans hours or days, the environment can be paused and resumed with its full working context intact, so the agent picks up precisely where it left off.
There is real engineering behind this. Each environment runs inside its own hardware-isolated microVM, which is what makes strong separation and sub-second startup possible at the same time. But the mechanism is not the point. Isolation is built into the runtime rather than wrapped around it, so teams stop choosing between a capable agent and a controlled one.
Nothing disappears into the sandbox. The agent is still governed by Foundry, so what it sent into the sandbox and what came back stay on the record with every other step it took.
The architecture pattern emerging across the enterprise
Put the two together and a clear design pattern appears, one we are now seeing repeatedly across industries.
Build and govern the agent on Foundry. Extend its execution into an isolated sandbox on Azure Container Apps. The agent keeps its identity, its permissions, and its oversight. The agent does not change when its work moves. Only the ground it runs on does. The work it performs happens in an environment that is purpose-built for agent tasks with ultra-fast executions, scoped to what it needs, and dissolved afterward.
This is what allows an organization to move from a handful of supervised agents to thousands running concurrently, without asking security and platform teams to accept risk they should never be asked to accept.
What this looks like in practice
Regulated client work, delivered at scale. Digital Gateway Powered by Claude brings Claude’s capabilities into KPMG’s connected tax platform on Microsoft Azure, giving professionals a secure, trusted environment to work alongside AI. Within the platform, DG Cowork serves as the AI-powered workspace where professionals can analyze information, generate content, and advance complex work. Because client data must remain separated by engagement, the platform was built with engagement-specific workspaces and controls at its foundation. DG Cowork is designed for a global scale, with more than 30,000 Azure Container Apps Sandboxes running concurrently.
Delivering answers to the questions nobody had time to ask. Cognite serves industrial operators who face a common challenge: their most valuable questions—like which assets are underperforming and why— are often too time-consuming to answer. Because traditional analysis took days, these investigations were rarely commissioned, leaving teams to rely on experience alone. Cognite Atlas AI changes that. Each agent gets its own governed workspace grounded in live customer data in Cognite Data Fusion, turning days-long investigations into traceable answers delivered in minutes. And for more complex problems, agents can pause and resume investigations without losing context, allowing them to work through a problem over time much like an engineer would.
Atlas AI agents are powered by Azure OpenAI models in Microsoft Foundry, but to handle complex work in Cognite Data Fusion they needed to execute arbitrary code and work with files directly – which meant sandboxes that truly isolate each user’s agent, environment, and data. With Azure Container Apps Sandboxes we had a prototype running in hours and were testing with customers within weeks, because Azure handles the hard part: per-user isolation, egress policies, sub-second execution. For our industrial customers, that’s the difference between a days-long manual investigation into “which wells are underperforming and why?” and a cited draft in minutes.
—Christian Flasshoff, Architect, Atlas AI, Cognite
A safe place for every learner. The Department for Education in South Australia runs EdChat so students can learn by writing code and exploring data alongside AI, across 60,000 students and more than 40,000 staff. That model only works if every student gets an environment of their own, with clear guardrails, and can come back later to find their work exactly as they left it. Building that in house meant owning the machinery behind it. The department estimates that moving to Azure Container Apps Sandboxes lets them retire close to 50,000 lines of code written to manage custom code interpreter state themselves. It is a per-user execution model at the scale of a school system, and a maintenance burden the department no longer has to carry.
Three very different organizations. One common story. None of them were held back by what the models could do. They were held back by not having anywhere to run the work at their scale, against their real data, with economics that made sense.
Learn more about Azure Container Apps Sandboxes
We built this because we needed it
More than a million sandboxes per day run in production across Microsoft, powering GitHub Copilot, Copilot Studio, Security Copilot, Foundry Agent Service, Azure SRE Agent, and other Azure services.
Every one of those sandboxes has someone on the other end of it. A developer expecting a pull request to come back finished. An analyst chasing an alert before the shift ends. A support team working down a queue that has to be empty by morning. Their day moves at the speed this layer runs, and it either holds or it doesn’t. That is a sterner test than anything we could design internally, and it is the one this infrastructure has been passing at scale for a long while now.
One architecture sits beneath both our developer tools and our AI platform. That is the same architecture we are making generally available to you.
The decision in front of you
Every major platform shift eventually comes down to a design choice made early, by people who could see where things were heading.
The teams who will scale agents successfully over the next few years are making that choice now. They are separating the platform that governs the agent from the environment that runs its work, and they are treating that separation as a foundational decision rather than something to retrofit once a pilot succeeds.
The agents are ready. The question worth asking is whether your platform is designed for software that acts.
Azure Container Apps Sandboxes
Learn more about designing agent-first platforms with Azure.
Get started
The post Designing agent-first platforms: What changes when agents do the work appeared first on Microsoft Azure Blog.
Quelle: Azure
This first article in our resilience series draws on a conversation with Mark Russinovich about how resilience is changing in the AI era and what it takes to continuously validate it at scale.
What surprises me most about resilience failures is how ordinary the drift is. A workload is deployed across availability zones, but a health probe still points to a single dependency. A database supports failover, but the application’s connection string is pinned to one region. Nothing looks broken. The architecture diagram still shows a resilient design even as the operational reality underneath it changes.
For years, resilience was something you set up once: configure disaster recovery, write a runbook, run the occasional failover test. That kept the lights on, but it treated resilience as a project with an end date rather than a property you maintain.
So, when an availability zone or a region has a bad day, the question is not whether a recovery plan exists on paper. It is whether resilience is still true today, and whether the team can prove it.
The dependency that breaks a workload is also changing. On Microsoft’s FY26 Q4 earnings call, Satya Nadella put it plainly: “For whatever reason, if a given model goes away, then you can’t be left high and dry. You need to be able to still continue your cyber operations.”
Traditional disaster recovery planning assumes the critical dependency is infrastructure. Increasingly it is an AI model, an inference endpoint, a retrieval pipeline, or a service operating under capacity constraints. A workload can be perfectly healthy from an infrastructure perspective and still fail its users because that dependency is unavailable, throttled, or economically impractical to run. That dependency rarely appears on the diagram at all.
There is a second shift underneath the first. An architecture diagram assumes a human drew it and a human will read it. Both of those assumptions are ending, and the dependencies it describes are no longer all deterministic. That changes what it means to know your estate is resilient.
This is the first in a series on how we are helping customers move to resilience that is designed in, measured as the estate changes, and improved over time. We have written before about how to design a resilient workload. This is a different problem: knowing whether hundreds of workloads still match their design today and being able to prove it. It is where a large share of our roadmap investment is now going.
Why resilience drifts
Resilience has always been a shared responsibility. We provide the availability zones, a secondary region of choice, and the replication primitives needed to support resiliency, and those do not drift. What drifts is the other half of the bargain: whether a given workload still uses them the way it was designed to, after a year of changes nobody flagged as risky. Change is where this concentrates. Across the industry, roughly 70 percent of cloud outages are related to change in some way—not dramatic failures, but ordinary modifications whose blast radius nobody re-evaluated.
That is why change discipline matters as much as design. Internally, a change rolls out to a canary region first, then a pilot region, with bake times where health signals are watched before it goes any further. That practice came out of an incident of our own, and it is the same discipline the Well-Architected Framework describes as safe deployment practices.
Learn more about the Azure Well-Architected Framework
Disaster recovery is reactive by design. Teams set recovery objectives when a project ships, stand up replication, and then move on, with little ongoing visibility into whether those objectives still hold as the workload changes. Resilience is designed once and rarely revisited. Across a growing estate of availability zones and regions, the distance between the resilience that was designed and the resilience that actually exists widens quietly, and it usually surfaces only during an incident.
const currentTheme =
localStorage.getItem(‘blogInABoxCurrentTheme’) ||
(window.matchMedia(‘(prefers-color-scheme: dark)’).matches ? ‘dark’ : ‘light’);
// Modify player theme based on localStorage value.
let options = {“autoplay”:false,”hideControls”:null,”language”:”en-us”,”loop”:false,”partnerName”:”cloud-blogs”,”poster”:”https://cdn-dynmedia-1.microsoft.com/is/image/microsoftcorp/1117657-resilience-drift_tbmnl_en-us?wid=1280″,”title”:”Resilience drift”,”sources”:[{“src”:”https://cdn-dynmedia-1.microsoft.com/is/content/microsoftcorp/1117657-resilience-drift-0x1080-6439k”,”type”:”video/mp4″,”quality”:”HQ”},{“src”:”https://cdn-dynmedia-1.microsoft.com/is/content/microsoftcorp/1117657-resilience-drift-0x720-3266k”,”type”:”video/mp4″,”quality”:”HD”},{“src”:”https://cdn-dynmedia-1.microsoft.com/is/content/microsoftcorp/1117657-resilience-drift-0x540-2160k”,”type”:”video/mp4″,”quality”:”SD”},{“src”:”https://cdn-dynmedia-1.microsoft.com/is/content/microsoftcorp/1117657-resilience-drift-0x360-958k”,”type”:”video/mp4″,”quality”:”LO”}],”ccFiles”:[{“url”:”https://azure.microsoft.com/en-us/blog/wp-json/bloginabox/v1/get-captions?url=https%3A%2F%2Fwww.microsoft.com%2Fcontent%2Fdam%2Fmicrosoft%2Fbade%2Fvideos%2Fproducts-and-services%2Fen-us%2Fazure%2F1117657-resilience-drift%2F1117657-resilience-drift_cc_en-us.ttml”,”locale”:”en-us”,”ccType”:”TTML”}],”downloadableFiles”:[{“url”:”https://cdn-dynmedia-1.microsoft.com/is/content/microsoftcorp/1117657-resilience-drift_transcript_en-us”,”locale”:”en-us”,”mediaType”:”transcript”}]};
if (currentTheme) {
options.playButtonTheme = currentTheme;
}
document.addEventListener(‘DOMContentLoaded’, () => {
ump(“ump-6ab4e1548fa53″, options);
});
We learned this the hard way. Mark tells the story of the storage change that came close to taking Azure down, and what it permanently changed about how we deploy.
What a diagram can’t tell you
A diagram is a claim about a system, made once, by someone reasoning about the system as they believed it to be. It is useful, and it is not evidence. There are four things it structurally cannot tell you.
Whether the goal is being met right now. A diagram has no timestamp. Health modeling does: described in the Well-Architected Framework and now available through health models in Azure Monitor, it represents an application as a hierarchy of its components and the signals underneath them, so health is expressed in terms the business recognizes rather than as a wall of resource-level metrics. Paired with service level indicators, it answers the only question that matters during an incident: is this application meeting its objective right now? The standard that matters is not our own. A service is only healthy if the customer thinks it is healthy—we can believe it is fine, but if the customer is not seeing a healthy service, we have a problem.
What “resilient” even means for this application. On a diagram, resilient is an adjective. A resiliency goal makes it a threshold you either meet or miss and defining it at the level of the application rather than resource by resource is what makes the answer meaningful.
Whether the failover path actually works. Every diagram draws the arrow. Only a test proves it.
The resources nobody drew. A diagram shows what someone remembered. Generated Infrastructure-as-Code covers every resource in the application. The difference between those two sets is where drift begins—and increasingly the reader on the other end is an agent rather than a person, working from the outcome you asked for instead of the picture you drew.
This is how we run Azure. Rather than asking each team to declare what healthy means for their service, we standardized on service level indicators, then applied machine learning to observed behavior so that healthy is defined by what the service actually does rather than by what someone assumed it would do.
This is the part I find most interesting. We wrote our documentation, our schemas, and our templates for people. Increasingly the thing reading them is an agent, working from the outcome you asked for rather than the picture you drew, and generating the resources itself. When the author and the reader are both machines, a diagram is no longer even the medium the decision is made in.
const currentTheme =
localStorage.getItem(‘blogInABoxCurrentTheme’) ||
(window.matchMedia(‘(prefers-color-scheme: dark)’).matches ? ‘dark’ : ‘light’);
// Modify player theme based on localStorage value.
let options = {“autoplay”:false,”hideControls”:null,”language”:”en-us”,”loop”:false,”partnerName”:”cloud-blogs”,”poster”:”https://cdn-dynmedia-1.microsoft.com/is/image/microsoftcorp/1117657-proving-resilience_tbmnl_en-us?wid=1280″,”title”:”Proving resilience”,”sources”:[{“src”:”https://cdn-dynmedia-1.microsoft.com/is/content/microsoftcorp/1117657-proving-resilience-0x1080-6439k”,”type”:”video/mp4″,”quality”:”HQ”},{“src”:”https://cdn-dynmedia-1.microsoft.com/is/content/microsoftcorp/1117657-proving-resilience-0x720-3266k”,”type”:”video/mp4″,”quality”:”HD”},{“src”:”https://cdn-dynmedia-1.microsoft.com/is/content/microsoftcorp/1117657-proving-resilience-0x540-2160k”,”type”:”video/mp4″,”quality”:”SD”},{“src”:”https://cdn-dynmedia-1.microsoft.com/is/content/microsoftcorp/1117657-proving-resilience-0x360-958k”,”type”:”video/mp4″,”quality”:”LO”}],”ccFiles”:[{“url”:”https://azure.microsoft.com/en-us/blog/wp-json/bloginabox/v1/get-captions?url=https%3A%2F%2Fwww.microsoft.com%2Fcontent%2Fdam%2Fmicrosoft%2Fbade%2Fvideos%2Fproducts-and-services%2Fen-us%2Fazure%2F1117657-proving-resilience%2F1117657-proving-resilience_cc_en-us.ttml”,”locale”:”en-us”,”ccType”:”TTML”}],”downloadableFiles”:[{“url”:”https://cdn-dynmedia-1.microsoft.com/is/content/microsoftcorp/1117657-proving-resilience_transcript_en-us”,”locale”:”en-us”,”mediaType”:”transcript”},{“url”:”https://cdn-dynmedia-1.microsoft.com/is/content/microsoftcorp/1117657-proving-resilience_audio_en-us”,”locale”:”en-us”,”mediaType”:”audio”}]};
if (currentTheme) {
options.playButtonTheme = currentTheme;
}
document.addEventListener(‘DOMContentLoaded’, () => {
ump(“ump-6ab4e1548ff32″, options);
});
Mark on why we break our own services on purpose, what a game day actually tests, and why every failure mode we simulate is one that has really happened.
When the dependency is probabilistic
A model that disappears is the obvious risk. The subtler one is a model that answers, differently each time. It is a newer source of drift, and it does not behave like the old ones. Ask the same question of a model twice and you can get two different answers. That makes correctness harder to define, and it makes change harder to reason about: if you change the prompt, change the model, or change the harness and the skills around it, you have changed the software, and it deserves the same discipline as any other change.
Most teams smoke-test instead. It looks fine, so it ships. Evaluation is the step that gets skipped, and it is the one that matters: measuring the system against what you actually consider valuable, not simply whether it responded.
The first question is whether the system needs to be probabilistic at all. Stay as deterministic as you can and use AI where it earns its place, not everywhere. And where you can wrap a non-deterministic system in a deterministic check, do it. If an agent is only supposed to update dependency versions, have a second system verify that versions are the only thing that changed. Where a deterministic check is not possible, use adversarial review: a second agent whose whole job is to find what is wrong with the first one’s work.
And accountability does not transfer. When someone deploys an agent, someone remains answerable for what it does.
That applies to us as well. Our own incident triage system uses language models to answer which service is responsible for an incident—work that used to mean waking people up to argue over logs. It is genuinely faster, and it is still trust but verify.
const currentTheme =
localStorage.getItem(‘blogInABoxCurrentTheme’) ||
(window.matchMedia(‘(prefers-color-scheme: dark)’).matches ? ‘dark’ : ‘light’);
// Modify player theme based on localStorage value.
let options = {“autoplay”:false,”hideControls”:null,”language”:”en-us”,”loop”:false,”partnerName”:”cloud-blogs”,”poster”:”https://cdn-dynmedia-1.microsoft.com/is/image/microsoftcorp/1117657-ai-uncertainty_tbmnl_en-us?wid=1280″,”title”:”AI uncertainty”,”sources”:[{“src”:”https://cdn-dynmedia-1.microsoft.com/is/content/microsoftcorp/1117657-ai-uncertainty-0x1080-6439k”,”type”:”video/mp4″,”quality”:”HQ”},{“src”:”https://cdn-dynmedia-1.microsoft.com/is/content/microsoftcorp/1117657-ai-uncertainty-0x720-3266k”,”type”:”video/mp4″,”quality”:”HD”},{“src”:”https://cdn-dynmedia-1.microsoft.com/is/content/microsoftcorp/1117657-ai-uncertainty-0x540-2160k”,”type”:”video/mp4″,”quality”:”SD”},{“src”:”https://cdn-dynmedia-1.microsoft.com/is/content/microsoftcorp/1117657-ai-uncertainty-0x360-958k”,”type”:”video/mp4″,”quality”:”LO”}],”ccFiles”:[{“url”:”https://azure.microsoft.com/en-us/blog/wp-json/bloginabox/v1/get-captions?url=https%3A%2F%2Fwww.microsoft.com%2Fcontent%2Fdam%2Fmicrosoft%2Fbade%2Fvideos%2Fproducts-and-services%2Fen-us%2Fazure%2F1117657-ai-uncertainty%2F1117657-ai-uncertainty_cc_en-us.ttml”,”locale”:”en-us”,”ccType”:”TTML”}],”downloadableFiles”:[{“url”:”https://cdn-dynmedia-1.microsoft.com/is/content/microsoftcorp/1117657-ai-uncertainty_transcript_en-us”,”locale”:”en-us”,”mediaType”:”transcript”},{“url”:”https://cdn-dynmedia-1.microsoft.com/is/content/microsoftcorp/1117657-ai-uncertainty_audio_en-us”,”locale”:”en-us”,”mediaType”:”audio”}]};
if (currentTheme) {
options.playButtonTheme = currentTheme;
}
document.addEventListener(‘DOMContentLoaded’, () => {
ump(“ump-6ab4e154902b2″, options);
});
Mark on why guardrails around AI should be deterministic wherever possible, and why the agent is never the one accountable.
What this looks like in practice
That discipline has to hold wherever a workload runs, and the foundations are familiar. The reliability guidance in the Well-Architected Framework already says to design resilience in from the start, and the Cloud Adoption Framework describes the operating half: carrying that intent into how the estate is actually run and reviewed month after month. Where an estate spans global, national, and sovereign or regulated environments, the word resilient has to mean the same thing in each rather than being redefined at every boundary; our guidance on reliability and sovereignty covers where those constraints intersect.
Your diagram shows three zones. It does not show that the health probe behind the load balancer resolves to one of them. Availability zones protect a workload from datacenter-level failures within a region, but only if compute, storage, and data tiers are genuinely spread across them.
Plan region resiliency for disaster recovery. Set explicit recovery objectives, an RTO and RPO, for each workload. Availability zones give you high availability inside a region, while a secondary region of choice provides a failover location when an entire region is affected. They answer different risks, and a resilient design is deliberate about both. In a regulated estate the choice narrows further, because a recovery region has to sit inside the same jurisdiction as the workload it protects.
Decide how much resilience is worth buying. Resilience is a cost decision as much as a design one. An application carrying 100 million dollars of revenue on a single day justifies an active-active topology across regions for that day; the same application may run in a single region with active-passive failover for the rest of the year. The right answer is deliberate, not maximal.
Know your blast radius. Understand what each workload depends on and where its hidden single points of failure are, then keep that picture current as the application changes. Internally we do this as reliability threat modeling: the same discipline as security threat modeling, asking of each component what would happen if it failed, and what we would do about it.
Check the dependencies your recovery path itself relies on. A workload can be replicated correctly and still be unrecoverable: if its encryption keys live only in the primary region, they are gone precisely when a region outage means you need them. Recovery paths have dependencies too, and they are rarely on the diagram.
Your diagram probably has no box for the AI dependency your application now relies on. Design for your application, workload, AI model, and service dependencies, not just infrastructure. Plan for graceful degradation and fallback so that if a critical dependency is deprecated, throttled, unavailable, or capacity-constrained, the workload continues to operate through an alternative path rather than failing outright.
This is not theoretical. Carne Group, one of Europe’s largest independent third-party asset managers with one trillion dollars under management, rebuilt its estate on Azure with infrastructure-as-code landing zones precisely so resilience would be reproducible rather than remembered. Because the definition lives in code, their team can stand up a duplicate site in another region and, as Carne Group’s global technology lead Stéphane Bebrone puts it, “even in the event of a worst-case scenario, we could be back up and running more or less in the same day.” They are working toward an active-passive topology across regions and plan to use Chaos Studio to verify those failover paths on a schedule rather than on an incident. Under DORA, they have to be able to prove it, not assert it.
Learn more about how Carne Group used Azure Site Recovery and Backup
Closing the gap between intent and reality
An architecture diagram is a statement of intent, not proof. It may show zone redundancy, regional failover, and protected dependencies, but only a test can determine whether those assumptions still hold. Assembling the underlying data into a clear view of application resiliency takes real effort and expertise. That is the customer’s half of the shared responsibility, and it is the gap Azure Infrastructure Resiliency Manager is built to help close. In public preview, its agentic-first experience helps teams start resilient, get resilient, and stay resilient.
Start resilient. Define what resilient means for an application, set explicit resiliency goals, and use the Resiliency Agent to generate resiliency-aware Infrastructure-as-Code up front, so a new workload starts resilient instead of being corrected months later. Service Groups help teams represent an application as a logical group of Azure resources in the portal, making it easier to manage resiliency posture at the application level rather than resource by resource.
Get resilient. See where the estate stands against the intended design, identify resources that were never zone resilient or stopped being so after a change, and prioritize the gaps that matter first. Rather than correlating findings across several tools, teams get recommendations and generated Infrastructure-as-Code for supported fixes, so remediation can move through a pull request instead of becoming a separate project.
Stay resilient. Validate the assumptions behind the recovery path before a real outage does it for you. For workloads where the customer manages the compute, such as virtual machines, a zone-down drill can simulate the loss of an availability zone and show what actually happens to the application. For other services, teams can use failover validation, product-specific recovery capabilities, and fault injection through Azure Chaos Studio to test the right failure modes.
We will be candid about the gaps too. Consistent, self-service resiliency assessment across every workload and environment is not finished work, and saying so matters. Resiliency improves when teams can see the gaps, measure them, and systematically close them over time. It is also work that never quite ends. As services harden the failure rate drops, but the failures that remain get rarer and stranger; you approach perfection asymptotically without ever arriving.
The bottom line: a diagram is not proof
Resilience is not a project you finish; it is a posture you maintain. The work is to design it in from the first architecture decision, then keep proving it as the estate changes, so drift is caught by a test rather than by an incident. All of it is anchored in the Azure Essentials frameworks, the Well-Architected Framework, and the Cloud Adoption Framework, so that resilient and sovereign carry the same meaning across every environment instead of being reinvented team by team.
None of this is a new framework. A diagram tells you what you intended; only a test tells you what you have. That was true when people drew the diagrams and read them. It matters more now that neither is reliably the case, and that some of what you depend on answers differently every time you ask. If resilience cannot be tested, it cannot be trusted.
In the next part of this series, we will look at how to measure that posture at scale.
const currentTheme =
localStorage.getItem(‘blogInABoxCurrentTheme’) ||
(window.matchMedia(‘(prefers-color-scheme: dark)’).matches ? ‘dark’ : ‘light’);
// Modify player theme based on localStorage value.
let options = {“autoplay”:false,”hideControls”:null,”language”:”en-us”,”loop”:false,”partnerName”:”cloud-blogs”,”poster”:”https://cdn-dynmedia-1.microsoft.com/is/image/microsoftcorp/1117657-full-interview_tbmnl_en-us?wid=1280″,”title”:”Full interview”,”sources”:[{“src”:”https://cdn-dynmedia-1.microsoft.com/is/content/microsoftcorp/1117657-full-interview-0x1080-6439k”,”type”:”video/mp4″,”quality”:”HQ”},{“src”:”https://cdn-dynmedia-1.microsoft.com/is/content/microsoftcorp/1117657-full-interview-0x720-3266k”,”type”:”video/mp4″,”quality”:”HD”},{“src”:”https://cdn-dynmedia-1.microsoft.com/is/content/microsoftcorp/1117657-full-interview-0x540-2160k”,”type”:”video/mp4″,”quality”:”SD”},{“src”:”https://cdn-dynmedia-1.microsoft.com/is/content/microsoftcorp/1117657-full-interview-0x360-958k”,”type”:”video/mp4″,”quality”:”LO”}],”ccFiles”:[{“url”:”https://azure.microsoft.com/en-us/blog/wp-json/bloginabox/v1/get-captions?url=https%3A%2F%2Fwww.microsoft.com%2Fcontent%2Fdam%2Fmicrosoft%2Fbade%2Fvideos%2Fproducts-and-services%2Fen-us%2Fazure%2F1117657-full-interview%2F1117657-full-interview_cc_en-us.ttml”,”locale”:”en-us”,”ccType”:”TTML”}],”downloadableFiles”:[{“url”:”https://cdn-dynmedia-1.microsoft.com/is/content/microsoftcorp/1117657-full-interview_transcript_en-us”,”locale”:”en-us”,”mediaType”:”transcript”},{“url”:”https://cdn-dynmedia-1.microsoft.com/is/content/microsoftcorp/1117657-full-interview_audio_en-us”,”locale”:”en-us”,”mediaType”:”audio”}]};
if (currentTheme) {
options.playButtonTheme = currentTheme;
}
document.addEventListener(‘DOMContentLoaded’, () => {
ump(“ump-6ab4e15490faf”, options);
});
The three clips above are drawn from a longer conversation covering the 2014 change that came close to taking Azure down, how we learned to measure health from the customer’s point of view, what breaks differently when a dependency is an AI model, and why the hardware in a datacenter fails every single day. Check out the full interview on the Azure Essentials YouTube channel.
Azure Essentials
Check out the full interview on the Azure Essentials YouTube channel.
Watch now
Resources
Cloud Resiliency
Azure Essentials
Azure Infrastructure Resiliency Manager
Reliability design principles—Azure Well-Architected Framework
Health modeling for workloads—Azure Well-Architected Framework
Health models in Azure Monitor (preview)
Cloud Adoption Framework for Azure
Azure Architecture Center
Reliability and Sovereignty in Azure
What’s new in the Resiliency service
Enable Zone Resiliency for Azure Workloads
Azure Chaos Studio documentation
Azure Site Recovery documentation
Azure Backup documentation
Carne Group boosts infrastructure and cyber resilience with Azure Site Recovery and Backup
Safe deployment practices—Azure Well-Architected Framework
Meet Brain: The AI system behind Azure reliability
Optimizing incident management with AIOps using the Triangle System
Proving application resilience on Azure with Chaos Studio
Advancing reliability—Azure Blog series
The post Your architecture diagram is not your resilience appeared first on Microsoft Azure Blog.
Quelle: Azure
Agents need containment, and a sandbox is only half of it. Something still has to say which agent runs there, what it gets, and what it may touch. That is a Kit: an ordinary OCI image, so the answer travels with the agent and means the same thing on any conforming runtime. Today we published the Docker Sandbox Kit Specification v3, open source under Apache 2.0 at docker/sandbox-kit-spec. Here is why I wrote it.
Everything that makes an agent useful is a grant
I run a lot of agents. They write code, run tests, install dependencies, call APIs, and work on infrastructure while I do something else. None of it happens without access, so I grant it one piece at a time: a bind mount, a token with broader scope than the task needs, a firewall rule that was quicker to open than to narrow. Each grant is reasonable on its own. Together they take back the isolation I was relying on, and none needed an exploit. The holes are configuration, added on purpose, usually by me.
I am worse at taking any of it back, and I could not reproduce the grants my setup depends on. No file records them. They live in shell history, dashboards, and my memory. I cannot hand that to a colleague or diff it against last week.
Containers package applications. Sandboxes contain agents.
A container packages applications. It shares the host kernel and uses namespaces and cgroups to give one fixed workload its own view of the filesystem, network, and processes. That is the right tool for software that runs, does its job, and touches only what it was handed.
An agent, however, is a probabilistic actor. It decides what to do next and then does it, to my filesystem, network, credentials, and cloud account. It will install a package that needs root, open a port nobody planned for, and try the next thing when the first is blocked. A container was not built for that: the boundary is the same kernel the actor is probing.
A Docker Sandbox is a microVM with its own kernel, so the boundary sits below anything the model can reach or rewrite. Inside one I can hand an agent root and let it loose, because the damage stops at the sandbox boundary. The sandbox is what lets me run an agent with the safeties off.
But an empty sandbox is not an environment. Something still has to say which agent runs, which tools and MCP servers it gets, which skills and instructions shape it, and exactly what it may touch.
What a Dockerfile cannot say
A Dockerfile answers everything about the software itself: how it is built, what gets packaged, how it starts. It was never standardized; OCI standardized the image it produces and how registries distribute it. What a Dockerfile does not describe is the outside: networks, credentials, volumes, tools, context. That half has lived in docker run flags, a Compose file, a CI config, and someone’s memory. Unversioned, unreviewable. A Kit writes it down with the content.
Where this came from
If you have used sbx, you have used Kits already. This is the third version of the format.The change that matters in this version is that a Kit is now an ordinary OCI image rather than its own artifact. Everything else in this post follows from that.
One image, one digest
If you have used sbx, you have used Kits. This is the third version of the format, and the change that matters is that a Kit is now an ordinary OCI image rather than its own artifact: no media type, no sidecar file, nothing for a registry to learn. The manifest carries the declarations in one annotation, vnd.docker.sandbox.kit.descriptor; the layers carry the content.
A Kit therefore builds with docker buildx build, pulls with docker pull, gets scanned and signed by the tooling you already run, and works in a FROM. Pinning the digest pins content, declarations, and metadata together. The tooling and distribution path are free; the format is something you learn: a grammar, a page per capability type, provides and requires, kind: set.
Two kinds of Kit exist. A workload runs and supplies the root filesystem. A mixin is an overlay: a CLI with its network rule, a credential binding, context for an agent. You launch one workload and any number of mixins.
Authority you can read
Part of the GitHub CLI mixin in the repository:
capabilities:
– type: com.docker.sandbox/network-policy@2
config:
runtime:
allow:
– github.com
– hosts: [api.github.com]
methods: [GET, HEAD, POST, PATCH, PUT, DELETE]
deny:
– hosts: [api.github.com]
methods: [DELETE]
paths: [/repos/**]
– type: com.docker.sandbox/credential@1
optional: true
config:
service: github
phase: runtime
apiKey:
name: GH_TOKEN
proxyManaged: true
inject:
– {domain: api.github.com, header: Authorization, format: "Bearer %s"}
Read it as a permission slip. This Kit asks to reach GitHub and nowhere else, and for most of the API but not deletes under /repos/**, because deny wins. The token that can open a pull request cannot delete the repository. The credential is proxy-managed: a conforming runtime injects the real value into requests to the named domains, and inside the sandbox there is only a sentinel.
Two words carry weight: asks and conforming. A Kit grants itself nothing. Each entry is a request, and the host decides. A conforming runtime, one that implements the behaviour the specification describes, blocks hosts not on the list. Without one, the annotation is inert: an image and no enforcement. Docker Sandboxes is the first conforming runtime.
Everything a Kit needs goes through that one list, typed and versioned. Grants (network rules, credentials, volumes, ports, devices, skills paths) count toward “what may this Kit do”; entries that ask the runtime to act, like a lifecycle hook, do not. A required request the host cannot satisfy refuses the launch, rather than starting an agent with less authority than it declared, or more.
Composition is a function, not a sequence
Container images never solved multiple inheritance: a Dockerfile stage has one FROM. Mixins are overlays ordered by the dependency graph the Kits declare through provides and requires, never by the order you typed the flags, so the same set always composes to the same image.
The resolver is strict on purpose. Every requires is satisfied from inside the set or resolution fails; nothing is fetched to cover a gap. Exactly one workload is allowed. Two Kits providing the same name fail rather than one silently shadowing the other (composing the Claude workload with the Claude mixin is the canonical mistake). Where Kits overlap, declarations reconcile: network rules union, hooks run in dependency order, guidance becomes one document, licenses union. Incompatible requests are an error, not a coin flip.
A kind: set descriptor names other Kits; publishing it runs the same coherence rules at build time and merges them into one ordinary Kit. An incoherent set fails at your build, not at someone else’s launch.
The diff is the review
The Claude Code Kit in the repository declares the hosts it asks to reach, its credential, the volumes that persist between sessions, and its install and startup hooks. When the next version asks for another host or a second credential, that is a change in authority, not a software update, and it shows up in the pull request as added lines a human can refuse.
Review depends on somebody reading the diff, so the specification defines a second gate that does not. Every descriptor reduces to a normalized set of everything the host would have to grant; a runtime that gates updates records that set and compares the next version against it. A version inside what was granted may apply without asking. Any widening stops and asks, and removing a deny rule counts: if a later gh Kit dropped DELETE /repos/**, the runtime holds the upgrade. That is why the declarations had to live in the artifact, not beside it.
Why this is a specification and not a feature
A Kit that stopped meaning anything when run somewhere else would be lock-in, not a trust boundary. So the grammar is normative, every capability type has its own page describing what a conforming runtime must implement, and types version independently (network-policy@1 and @2 both exist today). Two conformance suites ship with it: one judges whether an artifact is a conforming Kit, the other whether a runtime behaves as the pages say. Every normative statement is covered by a check or a written waiver.
Docker maintains the specification today, and it should not stay under a single vendor: a format for deciding what an agent may do is worth less if it belongs to whoever sells you the runtime. Docker Sandboxes will be a first-class implementation, not the only one. If a Kit you want cannot be expressed, or a runtime duty cannot be implemented as stated, open an issue.
Try it
sbx is our sandbox CLI (brew install docker/tap/sbx). From a checkout of the repository:
cd examples
sbx run ./hello –kit ./gh .
Edit a descriptor and only that Kit rebuilds; docker buildx build publishes it to any registry. Docker Cloud Sandboxes runs the same Kits with the same trust model on elastic capacity. The specification, capability pages, and a worked tour are in docker/sandbox-kit-spec.
Not only agents
Agents forced this into the open because the authority they ask for is so large, but ordinary workloads have always arrived with unwritten expectations: the endpoints they call, the credentials they need, the volume that must survive a restart. That knowledge has lived in a Helm chart, a runbook, or a colleague. It is the same gap, less alarming when a web service gets it wrong. This specification is where any software writes down what it needs from the world around it; agents were the case urgent enough to have it built.
Dockerfiles made software reproducible. Kits make authority reproducible.
Quelle: https://blog.docker.com/feed/