GPT-6 Astra: Frontier intelligence for work, now generally available in Microsoft Foundry

The next era of enterprise AI will not be defined by chat experiences. It will be defined by how well a model can work for and with you. GPT-6 Astra, OpenAI’s newest frontier model, is now generally available for all customers in Microsoft Foundry. It is designed to help organizations make decisions for complex work and execute across the applications and systems where your business operates.

We are investing to make it easier for customers to use advanced technology like Astra. AI initiatives often slow down on identity, networking, governance, data handling, evaluation, and compliance. Microsoft Foundry brings these fundamentals together in Azure, helping teams move from experimentation to production with speed and trust.

Turn open-ended goals into action

Astra is built to take an open-ended challenge, reason through it in multiple steps, create a plan, and produce a polished result. It can weigh trade-offs, incorporate new direction as work progresses, and use tools across applications and systems. 

For enterprises, this shifts AI from conversational assistance toward delivering more substantial units of work:

Deliberate planning and decision support. Astra can break a challenge into steps, evaluate options, communicate its recommendation, and identify the next actions for review.

Polished, purposeful output. Astra can apply context, templates, and quality standards throughout a workflow, helping produce documents, spreadsheets, presentations, and analyses that are ready for review.

Execution across applications. With advanced tool use and computer use, Astra can interact with software on a person’s behalf, move between apps, and complete multi-step tasks with appropriate human oversight.

At Replit, our mission is making useful intelligence accessible to everyone. GPT-6 Astra available through Microsoft Foundry unlocks a new level of agentic capability that goes beyond code generation to active software creation and more. We’re excited about the opportunities created for developers and entrepreneurs to build more ambitious applications with an intelligent software-building partner.
—Luis Hector Chavez, CTO, Replit

Computer use across applications

Astra’s computer-use capabilities are designed to work across familiar applications, including workflows without dedicated APIs. It can interpret on-screen information and interact with approved interfaces to support tasks such as updating records, navigating development tools, testing software, and assembling results into reports. OpenAI reports state-of-the-art results on selected computer-use evaluations; performance varies by task, tools, configuration, and safeguards.

Capability this direct demands containment. Content displayed in an application may be incomplete, misleading, or designed to influence an agent’s behavior. Foundry helps customers define access, approvals, and monitoring, and design workflows with scoped credentials, approved resources, human checkpoints for consequential actions, and activity records aligned to their risk requirements.

Enterprise scenarios we’re seeing

Software engineering: Astra can reproduce complex bugs, investigate likely causes, propose fixes, and prepare changes for developer testing and review.

Business intelligence: Astra can build and refine dashboards in Power BI, helping analysts compare data, identify trade-offs, and prepare insights to share.

Professional work: Astra can produce documents, spreadsheets, and presentations that follow existing templates and business standards, creating polished artifacts for expert review.

Application workflows: Astra can support tasks such as updating customer records, processing forms, testing websites, and working through approved interfaces where dedicated APIs are limited. 

Enterprise controls for agentic work

OpenAI describes Astra as its most aligned model to date and plans to publish supporting alignment, safety, and computer-use evaluations in its supporting launch materials. Foundry complements that model-level work with enterprise security, safety, and compliance capabilities, including Microsoft Entra identity and access management, encryption in transit and at rest, private networking options, role-based access controls, content filtering, safety evaluations, monitoring, and governance tools.

Prompts and outputs are not used to train the models. These capabilities help customers configure safeguards and maintain oversight, but do not eliminate risk or replace each organization’s responsibility to select and configure controls appropriate to its scenarios and regulatory obligations.

At Albertsons Companies, we believe the real advantage in frontier AI is the ability to evolve as quickly as the technology does, without compromising enterprise discipline. That means creating an environment where we can evaluate new capabilities, put the right ones to work quickly and maintain consistent security, governance and operational controls as we scale. Azure OpenAI on Microsoft Foundry helps us create that balance of speed and control, so our teams can stay focused on delivering meaningful outcomes for our customers, associates and the business.
—Anirban Nandi, VP, Data and AI, Albertsons Companies

Global scale with service level to match

GPT-6 Astra is available in Foundry models with Standard and Provisioned Throughput deployment options, in both Global and US Data Zone geographies. Standard provides pay-as-you-go flexibility for variable demand, while Provisioned Throughput provides dedicated model-processing capacity for workloads requiring consistent latency and guaranteed throughput. Customers can choose the right deployment based on workload requirements.

Astra is designed for token efficiency on complex work, helping customers manage consumption as they scale. Actual usage and costs will vary by workload and configuration.

GPT-6 Astra pricing*

Deployment Context Length Pricing (USD $/million tokens) Input Cached Input Cached Writes Output Standard Global Short context $10.00 $1.00 $12.50 $50.00 Long context $20.00 $2.00 $25.00 $75.00 Standard Data Zone (US) Short context $11.00 $1.10 $13.75 $55.00 Long context $22.00 $2.20 $27.50 $82.50 Provisioned Throughput pricing varies by deployment type. U.S. Data Zone Provisioned Throughput is priced at a 10% premium to Global Provisioned Throughput. For current rates and terms, see the Azure OpenAI pricing page.

Get started today

Explore the model: Try GPT-6 Astra in Foundry Models to see its advanced capabilities firsthand.

Build agentic workflows: Start with the Foundry Agent Service to bring cross-application task execution to your workflows.

Try GPT-6 Astra today

Explore the model and start building agentic workflows.

Learn more

The post GPT-6 Astra: Frontier intelligence for work, now generally available in Microsoft Foundry appeared first on Microsoft Azure Blog.
Quelle: Azure

How Microsoft’s Physical Security Engineering Team scaled hybrid operations with Azure Arc and Azure Virtual Desktop

When a physical security operator begins a shift supporting Microsoft’s global datacenter operations, they depend on a collection of applications and systems that help monitor access activity, review video feeds, investigate alerts, and coordinate physical security operations across a complex global environment. Those tools must be available, responsive, and reliable from the moment a shift begins.

As Azure datacenters expanded to support growing demand for cloud and AI services, maintaining that experience became increasingly important. Critical security systems were distributed across hundreds of locations worldwide, while the infrastructure supporting them spanned both on-premises and cloud environments. The challenge wasn’t responding to a specific incident or operational failure but ensuring that as Azure’s physical footprint continued to grow, the systems supporting those operations remained secure, manageable, observable, and consistent at a global scale.

See how Azure Arc unifies hybrid operations

Meeting that goal required more than simply keeping systems online. The team needed a way to manage infrastructure across hybrid environments, standardize operations, automate routine tasks, improve visibility into system health, and provide operators with consistent application experiences regardless of location. By combining Azure Arc, Azure Virtual Desktop, Azure Monitor, and other Azure management services, Microsoft built a more unified operational foundation designed to support the evolving needs of its global physical security environment.

Building a unified management layer across hybrid infrastructure

As Azure datacenters expanded, so did the infrastructure supporting their physical security operations. Critical systems were deployed close to the environments they served and operated within highly segmented networks designed to prioritize resiliency, security, compliance, and local autonomy. That architecture solved one challenge but created another.

The physical security organization was responsible for deploying and managing thousands of servers distributed across Microsoft’s global datacenter footprint in alignment with established protocols. While each deployment met baseline operational requirements, rapid growth and increasing scale made it increasingly difficult to guarantee consistency.

The team needed a way to bring these distributed systems under a common management framework without changing where the workloads ran or weakening the security boundaries that protected them.

Why Azure Arc

The objective wasn’t to move these workloads into Azure. Many of the systems supporting physical security operations needed to remain close to the environments they served and continue functioning independently when required by local operational or resiliency needs. Instead, the team was looking for a way to extend the operational benefits of Azure to on-premises infrastructure.

Azure Arc was designed to address exactly this type of challenge. At its core, Azure Arc extends Azure’s management and governance capabilities to servers and resources running outside Azure. Rather than treating on-premises systems as separate operational islands with their own tools and processes, Azure Arc allows organizations to manage those resources through Azure’s control plane. This makes it possible to apply, at scale, many of the same monitoring, policy, automation, security, and update-management workflows used in Azure to infrastructure running elsewhere.

For Microsoft’s physical security organization, Azure Arc made it possible to manage servers across its global datacenter footprint through a common operational model, regardless of where they were physically located.

More importantly, Azure Arc allowed the team to preserve the resiliency and security characteristics of their existing deployments while gaining centralized visibility, governance, and automation capabilities.

Establishing a consistent operational foundation

Once onboarded to Azure Arc, the team began extending familiar Azure management capabilities to infrastructure running outside Azure. Using Azure Update Manager, patching activities that had historically required significant coordination across distributed environments could be scheduled, tracked, and governed through a centralized framework. According to the team, this automation now saves thousands of hours annually while enabling a relatively small operations team to support a growing infrastructure footprint.

At the same time, Azure Policy, Guest Configuration, Azure Monitor, Azure Monitor Agent, and Log Analytics helped create a common framework for governance, compliance monitoring, and observability. The team could continuously assess critical security configurations, identify drift, monitor system health, and surface operational telemetry through centralized dashboards, alerts, and reporting workflows regardless of where infrastructure was deployed.

Security remained a primary consideration throughout the design. Managed Identities and Azure role-based access control (RBAC) helped reduce reliance on stored credentials while providing more granular control over access to operational resources. Azure Automation further reduced manual effort by standardizing remediation, maintenance, and configuration-management activities through reusable runbooks. Together, these capabilities helped establish a more consistent operating model across the environment while improving visibility, strengthening governance, and reducing the operational overhead associated with managing a globally distributed infrastructure.

Delivering a consistent operator experience with Azure Virtual Desktop

Unified management solved one part of the challenge. The next was ensuring that operators interacting with those systems received the same level of consistency, performance, and visibility.

The team’s objective extended beyond providing remote access. They needed a way to improve application performance, simplify lifecycle management, and gain better insight into the end-user experience. Azure Virtual Desktop provided a flexible platform for delivering applications closer to the infrastructure they depended on, while also enabling centralized image management and integration with Azure monitoring services. This allowed the team to maintain consistent host configurations, simplify updates, and incorporate user-session telemetry into existing operational workflows.

To improve the operator experience, the team relocated the application environment closer to the infrastructure it supported and delivered access through Azure Virtual Desktop sessions. The impact was immediate: application launch times improved by approximately 12x, helping operators access critical tools more quickly and consistently.

The team also adopted a centralized image-management strategy and automated host refresh process. Instead of maintaining individual systems over time, hosts could be rebuilt from approved images and deployed consistently across the environment. This approach accelerated release cycles by ~6x, reduced configuration drift, and allowed updates that once required weeks or months of coordination to be completed in hours.

Equally important was the visibility Azure Virtual Desktop unlocked. By integrating Azure Virtual Desktop with Azure Monitor, Azure Monitor Agent, Log Analytics, and Azure Virtual Desktop Insights, the team gained access to telemetry on session health, round-trip time, bandwidth usage, and client-side application behavior. Engineers could better understand how applications performed from the operator’s perspective, identify trends earlier, and shift from reactive troubleshooting to a more proactive, data-informed approach.

Key lessons for managing hybrid environments at scale

As Azure’s global datacenter footprint continued to grow, Microsoft’s physical security organization needed a management and delivery model that could scale alongside it. By combining Azure Arc and Azure Virtual Desktop, the team established a more consistent approach to managing infrastructure, delivering applications, and monitoring operational health across a complex hybrid environment.

The result wasn’t a single breakthrough technology, but a unified operating model that improved visibility, reduced operational overhead, and helped ensure critical systems remained resilient, manageable, and ready to support future growth.

Learn more

Azure Arc

Azure Virtual Desktop

Azure Monitor

Bring consistency and control to hybrid operations

See how Azure Arc helps organizations extend Azure management and governance capabilities across distributed infrastructure, enabling centralized visibility, automation, compliance, and operational consistency without changing where workloads run.

Explore Azure Arc

The post How Microsoft’s Physical Security Engineering Team scaled hybrid operations with Azure Arc and Azure Virtual Desktop appeared first on Microsoft Azure Blog.
Quelle: Azure

GPT-6 Astra: Frontier intelligence for work, now available in Microsoft Foundry

The next era of enterprise AI will not be defined by chat experiences. It will be defined by how well a model can work for and with you. GPT-6 Astra, OpenAI’s newest frontier model, begins rolling out today through the Microsoft Foundry Limited Access Program, with availability expanding to participating customers over the coming days. It is designed to help organizations make decisions for complex work and execute across the applications and systems where your business operates.

We are investing to make it easier for customers to use advanced technology like Astra. AI initiatives often slow down on identity, networking, governance, data handling, evaluation, and compliance. Microsoft Foundry brings these fundamentals together in Azure, helping teams move from experimentation to production with speed and trust.

Turn open-ended goals into action

Astra is built to take an open-ended challenge, reason through it in multiple steps, create a plan, and produce a polished result. It can weigh trade-offs, incorporate new direction as work progresses, and use tools across applications and systems. 

For enterprises, this shifts AI from conversational assistance toward delivering more substantial units of work:

Deliberate planning and decision support. Astra can break a challenge into steps, evaluate options, communicate its recommendation, and identify the next actions for review.

Polished, purposeful output. Astra can apply context, templates, and quality standards throughout a workflow, helping produce documents, spreadsheets, presentations, and analyses that are ready for review.

Execution across applications. With advanced tool use and computer use, Astra can interact with software on a person’s behalf, move between apps, and complete multi-step tasks with appropriate human oversight.

At Replit, our mission is making useful intelligence accessible to everyone. GPT-6 Astra available through Microsoft Foundry unlocks a new level of agentic capability that goes beyond code generation to active software creation and more. We’re excited about the opportunities created for developers and entrepreneurs to build more ambitious applications with an intelligent software-building partner.
—Luis Hector Chavez, CTO, Replit

Computer use across applications

Astra’s computer-use capabilities are designed to work across familiar applications, including workflows without dedicated APIs. It can interpret on-screen information and interact with approved interfaces to support tasks such as updating records, navigating development tools, testing software, and assembling results into reports. OpenAI reports state-of-the-art results on selected computer-use evaluations; performance varies by task, tools, configuration, and safeguards.

Capability this direct demands containment. Content displayed in an application may be incomplete, misleading, or designed to influence an agent’s behavior. Foundry helps customers define access, approvals, and monitoring, and design workflows with scoped credentials, approved resources, human checkpoints for consequential actions, and activity records aligned to their risk requirements.

Enterprise scenarios we’re seeing

Software engineering: Astra can reproduce complex bugs, investigate likely causes, propose fixes, and prepare changes for developer testing and review.

Business intelligence: Astra can build and refine dashboards in Power BI, helping analysts compare data, identify trade-offs, and prepare insights to share.

Professional work: Astra can produce documents, spreadsheets, and presentations that follow existing templates and business standards, creating polished artifacts for expert review.

Application workflows: Astra can support tasks such as updating customer records, processing forms, testing websites, and working through approved interfaces where dedicated APIs are limited. 

Enterprise controls for agentic work

OpenAI describes Astra as its most aligned model to date and plans to publish supporting alignment, safety, and computer-use evaluations in its supporting launch materials. Foundry complements that model-level work with enterprise security, safety, and compliance capabilities, including Microsoft Entra identity and access management, encryption in transit and at rest, private networking options, role-based access controls, content filtering, safety evaluations, monitoring, and governance tools.

Prompts and outputs are not used to train the models. These capabilities help customers configure safeguards and maintain oversight, but do not eliminate risk or replace each organization’s responsibility to select and configure controls appropriate to its scenarios and regulatory obligations.

At Albertsons Companies, we believe the real advantage in frontier AI is the ability to evolve as quickly as the technology does, without compromising enterprise discipline. That means creating an environment where we can evaluate new capabilities, put the right ones to work quickly and maintain consistent security, governance and operational controls as we scale. Azure OpenAI on Microsoft Foundry helps us create that balance of speed and control, so our teams can stay focused on delivering meaningful outcomes for our customers, associates and the business.
—Anirban Nandi, VP, Data and AI, Albertsons Companies

Global scale with service level to match

GPT-6 Astra will be available through Standard deployments including Global and U.S. Data Zone. Customers can choose the right deployment based on workload requirements.

This consumption-based model lets teams begin building without committing to reserved capacity. Astra is also designed for token efficiency on complex work, helping customers manage consumption as they scale. Actual usage and costs will vary by workload and configuration.

GPT-6 Astra pricing

DeploymentPricing (USD $/million tokens)InputCached InputCached WritesOutputStandard Global (Short context)$10.00$1.00$12.50$50.00Standard Global (Long context)$20.00$2.00$25.00$75.00Standard Data Zone (US) (Short context)$11.00$1.10$13.75$55.00Standard Data Zone (US) (Long context)$22.00$2.20$27.50$82.50

Get started today

Explore the model: Try GPT-6 Astra in Foundry Models to see its advanced capabilities firsthand.

Build agentic workflows: Start with the Foundry Agent Service to bring cross-application task execution to your workflows.

Try GPT-6 Astra today

Explore the model and start building agentic workflows.

Learn more

The post GPT-6 Astra: Frontier intelligence for work, now available in Microsoft Foundry appeared first on Microsoft Azure Blog.
Quelle: Azure

Enterprise AI transformation relies on the end-to-end platform: Azure was built for this moment

Summary
The recognition for Microsoft over the past couple of weeks comes down to models, infrastructure, data, applications, and developer tools working as one system when AI moves into production.

Enterprise AI is moving into production, and our customers are becoming multi-model. Organizations will use frontier models where capability matters, and smaller, specialized, and open-weight models where economics and finer controls matter. But the value does not come from any model in isolation. It comes from the system around it: infrastructure, data, applications, agents, security, and operations working together. That compounding value is what Microsoft Azure is built to deliver.

A system built from silicon to agent

That integration extends into the infrastructure underneath the model. Customers want the flexibility to choose across models and infrastructure without having to stitch together and tune every layer themselves. Microsoft has drawn on decades of running mission-critical systems and operating some of the world’s most demanding AI services at global scale. We believe that breadth and integration across the platform, extending through developer tools and AI applications is a key reason why Microsoft has been named a Leader in both the 2026 Gartner® Magic Quadrant™ for Strategic Cloud Platform Services and The Forrester Wave™: Public Cloud Platforms, Q3 2026.

We appreciate the recognition. What matters more is that customers choosing a platform today are shaping their infrastructure for years, and that choice rests on system-level capability. A cloud platform now must do more than provide individual services. It must give customers choice across models and infrastructure while helping them build faster, run reliably, manage risk, control cost, and improve outcomes. For an enterprise building the next generation of AI applications, how the layers work together matters more than any single feature.

Discover trusted cloud solutions on Microsoft Azure

Microsoft’s Leader placement in the 2026 Gartner® Magic Quadrant™ for Strategic Cloud Platform Services follows Leader placements in the 2025, 2024, and 2023 editions. We believe that what matters for customers, is whether the platform can translate technology into real impact: better performance, greater cost efficiency, faster delivery, and the ability to scale critical systems with confidence.

Microsoft was also named a Leader in The Forrester Wave™: Public Cloud Platforms, Q3 2026. Forrester’s evaluation looks at both the strength of the current offering and the strategy behind it. This recognition provides another independent view of how Azure is evolving as customers move from isolated AI projects to production systems.

Forrester describes Microsoft’s direction as a vision of Azure as a single, vertically integrated system.

Choice without complexity

A multi-model strategy does not mean every model should run the same way. The platform must support those choices across heterogeneous compute while applying consistent security, identity, governance, reliability, and operations.

Microsoft Foundry is central to this approach. It gives developers broad model choice and the tools to evaluate, secure, monitor, and operate AI systems, with Azure infrastructure underneath. This is not about forcing every workload into one model. It is about using reducing the seams between layers so teams can make workload-specific choices while operating consistently across cloud, on-premises, edge, and third-party environments.

Data gives AI its business value

Model choice will keep changing, but the data and business context that make AI useful endure. Customers want to work with data where it already resides, without creating more copies or losing governance along the way. As Forrester puts it: “Models come and go; data has gravity.”

Microsoft Fabric brings analytics and data together, and Microsoft Purview applies governance across that estate. The Azure databases, including Azure SQL and Azure Cosmos DB, connect AI to current operational data. On top of that foundation, Microsoft IQ provides the unified enterprise intelligence layer, giving apps and agents consistent business context across work, data, and knowledge. Together, these capabilities let organizations change models without rebuilding the data, governance, and business context around every application.

UNC Health illustrates why that foundation matters. By modernizing its analytics, the organization is creating a governed data environment that supports care, operations, and research within the requirements of a highly regulated industry. It is the kind of foundation organizations need before AI can be applied responsibly at scale.

Modernization is the catalyst to AI

The applications running a business today contain years of business logic, data, and operating knowledge. They need a modern home where they can continue to support proven processes and connect to new AI experiences. Modernization is therefore part of the AI work, not a separate project. Customers need to decide workload by workload whether to move it, update it, use a managed service, expose it to agents through secure interfaces, or rebuild the parts where there is a clear business reason.

Levi Strauss & Co. shows how modernization and AI become part of the same journey. The company modernized its legacy infrastructure on Azure to build a more resilient foundation, then used Microsoft Foundry to introduce agents that simplify work and accelerate decision-making. A heritage company did not have to leave its existing business behind to adopt AI; it modernized that foundation and built forward from it.

Agents can help teams assess applications, plan upgrades, refactor code, test changes, and support migration while developers and IT teams retain control of architecture and business decisions. GitHub Copilot agentic modernization supports .NET and Java applications, and the work connects across the software lifecycle. This is where the analyst feedback is especially relevant: Gartner highlights Microsoft’s pragmatic approach to application modernization and its integrated, end-to-end software developer lifecycle, while the Forrester report notes our customers’ appreciation for Microsoft’s migration and modernization expertise. The goal is straightforward: help customers modernize the applications they already rely on and so they are ready for the next generation of AI.

Power every AI ambition

As customers run more AI in production, the platform must be more efficient, more reliable, and easier to operate. Customers need the freedom to choose the models and infrastructure that best fit each workload, while the platform reduces the complexity of bringing those choices together.

We’re proud to be recognized as a Leader by Gartner and Forrester, and even more excited by what these evaluations reflect about where the industry is heading. We believe the next generation of cloud will be defined by how well the platform brings infrastructure, data, models, applications, and developer tools together while preserving the choice customers need as each layer continues to evolve.

That’s the direction we’re building toward with Azure, and we’re excited to keep shaping what comes next alongside our customers and partners.

Review the 2026 Gartner® Magic Quadrant™ for Strategic Cloud Platform Services. Read the report.

Review The Forrester Wave™: Public Cloud Platforms, Q3 2026. Read the report.

Gartner® Magic Quadrant™ for Strategic Cloud Platform Services, 2026. By Alessandro Galimberti, Carolin Zhou, Douglas Toombs, Dennis Smith, Ed Anderson, Tobi Bet, Chuck Lawton, 1 September 2026.

Gartner does not endorse any company, vendor, product or service depicted in its publications, and does not advise technology users to select only those vendors with the highest ratings or other designation. Gartner publications consist of the opinions of Gartner’s business and technology insights organization and should not be construed as statements of fact. Gartner disclaims all warranties, expressed or implied, with respect to this publication, including any warranties of merchantability or fitness for a particular purpose.

Gartner and Magic Quadrant are trademarks of Gartner, Inc., and/or its affiliates.

This graphic was published by Gartner, Inc. as part of a larger research document and should be evaluated in the context of the entire document. The Gartner document is available upon request here.

Forrester does not endorse any company, product, brand, or service included in its research publications and does not advise any person to select the products or services of any company or brand based on the ratings included in such publications. Information is based on the best available resources. Opinions reflect judgment at the time and are subject to change. This report is part of a broader collection of Forrester resources, including interactive models, frameworks, tools, data, and access to analyst guidance. For more information, read about Forrester’s objectivity here.
The post Enterprise AI transformation relies on the end-to-end platform: Azure was built for this moment appeared first on Microsoft Azure Blog.
Quelle: Azure

The Economics of Agent Optimization: Context engineering for enterprise AI agents

This blog post is the third of a four-part series called The Economics of Agent Optimization, which shares the strategies, capabilities, and proof points to help you optimize agent costs and run AI as a managed investment system on Microsoft Foundry. The first post set out the three decisions that systems rest on. The second post took the request at runtime. This post takes the next one: making each agent cheaper over time as it learns what works.

Every agent has a mechanism that determines what its model sees on each turn. In many production systems, that choice was set during prototyping and never revisited, even though it often drives the largest share of operating cost and contributes to disappointing answers.

This is also the part of an agent that can improve on its own. The model remains as capable as when you selected it, and instructions change only when someone rewrites them. But what an agent knows, can access, and remembers, grows as it runs—making it the key to improving performance while lowering cost over time. Managing that process is called context engineering.

Why the context window sets what an agent costs

A model has no memory of its own. On each turn, its context window supplies everything it can use: instructions, available tools, retrieved documents, and conversation history. When the turn ends, that context disappears and must be sent again on the next one.

That cost is manageable for a chatbot answering one question. For an agent working across many turns toward one outcome, it is often the largest expense. Because the context window is paid for every turn, unnecessary content is billed repeatedly.

The less visible cost is quality. More context does not guarantee better answers: a relevant fact buried in 40 pages is harder to use, and a long tool list makes the wrong choice more likely. Each mistake adds more turns—and more cost—to recover.

That makes context worth a leader’s attention. Most cost reductions involve a tradeoff: a cheaper model may reduce quality, and shorter instructions may weaken an answer. By contrast, removing unnecessary context can lower costs without reducing quality, making it an easier optimization for teams to support.

What context engineering means in practice

That is what context engineering does: it decides what enters the context window on each turn, so the agent gets what this request needs rather than everything it might ever need. As a one-time choice, it is a design decision. Practiced continuously, it is how an agent improves, because every turn reveals what it actually used. Four questions cover the work, and teams usually take them in this order.

What should the agent know?

Many teams begin with broad searches that insert entire documents into the prompt. This approach is easy to build but costly to run, and it forces the model to find the one relevant detail amid everything else.

Foundry IQ replaces that with a managed knowledge layer. A knowledge base points at sources across Work IQ, Fabric IQ, Web IQ, Microsoft Azure Blob Storage, SharePoint, OneLake, and Azure SQL. When an agent submits a query, Foundry IQ decomposes it into subqueries, searches connected sources in parallel, semantically reranks the results, and returns grounded passages with citations. This narrows what enters the model’s context to the most relevant evidence while preserving traceability to the source.

Two features make this knowledge layer reusable across agents and governable at scale. A single knowledge base can serve multiple agents. Indexed sources can refresh incrementally on a configured indexer schedule, while remote sources are queried on demand. At query time, Foundry IQ can run under the caller’s Microsoft Entra identity, synchronize access-control lists for supported sources, and honor Microsoft Purview sensitivity labels, so the agent retrieves only content the caller is authorized to access.

Our internal evaluations showed that Foundry IQ knowledge bases improved evidence recall by up to 54% on the BrowseComp-Plus benchmark while reducing retrieval token costs by 34%. The gains came from agentic retrieval, semantic reranking, improved answer synthesis, and more efficient token use.

What should the agent be able to reach?

Tool overhead is easy to miss: adding one may take a single line of code, but its full description occupies the prompt. Every tool attached to an agent has that description sent to the model on every turn, needed or not, and enterprise agents pick up tools quickly as they connect to more systems.

Toolboxes in Foundry give an agent one managed Model Context Protocol (MCP) endpoint for built-in tools like web search, code interpreter, and file search alongside custom MCP servers, OpenAPI 3.0 and 3.1 APIs, and A2A agents. Foundry manages authentication, access policies, and tool versions in one place, rather than configuring each integration separately for every agent. Once a new toolbox version is tested and promoted, connected agents can use it without code changes or redeployment.

Toolboxes organize your tools. The tool search capability inside Toolbox is what stops you paying for all of them. Instead of the full list, the model gets two things: a way to describe what it needs in plain language, and a way to call whatever comes back. The cost of the tool list stays flat, however large the toolbox grows. In internal benchmarking against a public, open-source tool-retrieval dataset, Toolboxes in Foundry reduced average input-token consumption around 97% for large tool libraries—directly lowering inference costs for customers building agents.1

Foundry also notices which tools each toolbox uses most and puts those within easy reach, so the common path gets faster and cheaper the longer the agent runs. Accuracy improves alongside cost, because a short, well-matched list means fewer wrong calls and fewer turns spent recovering.

How should the agent do the work?

Knowledge and tools cover what an agent can find and do. Neither covers how your company expects the work to be done: the escalation path a support agent follows; the checklist a code review applies. That guidance usually lives in the agent’s instructions. As a result, the same procedures may be copied across multiple agents and included in every request, even when they are not relevant.

A skill turns that guidance into a named, reusable procedure. Skills are stored centrally in Foundry and made available to agents through a toolbox. Instead of embedding a copy of the procedure in each agent, the toolbox references the centrally managed skill. When your organization improves a procedure, you can publish a new version and set it as the default. Every agent using that skill can then follow the updated procedure without code changes or redeployment. To minimize context usage, the agent initially sees only each skill’s name and short description. It loads the full instructions only when the skill is relevant. This makes it practical to offer a large library of detailed procedures without adding unnecessary content to every interaction.

What should the agent remember?

Agents need continuity, but they do not need to carry every detail from every interaction. Repeatedly sending an entire conversation to the model adds cost and consumes context, even when only a few details remain useful.

Memory in Foundry Agent Service helps agents retain important context without replaying entire conversations. It supports three types of memory:

Session memory for the current conversation.

User memory for preferences and facts that persist across sessions.

Procedural memory for learned workflows and task execution patterns.

This allows a returning customer to pick up where they left off, while enabling an agent to consistently follow proven processes without being re-instructed each time.

Together, these capabilities help an agent continue a customer interaction, personalize future responses, and improve how reliably it completes recurring tasks. Procedural memory complements centrally managed skills: a skill defines the organization’s approved procedure, while procedural memory helps an agent learn from its own task execution. In Microsoft’s evaluations, enabling procedural memory produced about a 5% improvement on STATE-Bench and Tau-Bench. Organizations can also control memory through user-level isolation, retention settings, and time-to-live policies that determine what is stored and when it expires.

Why context engineering becomes a system

Any team can assemble knowledge retrieval, tools, procedural guidance, and memory. The challenge is making them work together, under one set of permissions, and keeping them current as the organization changes.

Foundry brings these pieces into a single system. Knowledge, tools, skills, and memory can be managed through shared infrastructure rather than separate products, while permissions are enforced where data is retrieved, so agents inherit the access controls already applied to enterprise content. Foundry IQ extends that model across enterprise knowledge, business data, and organizational context, while remaining compatible with frameworks such as Microsoft Agent Framework, LangGraph, GitHub Copilot SDK, and Claude Agent SDK.

The result is that context improves without requiring agents to be rebuilt. Knowledge bases refresh as source systems change. Skills evolve as policies evolve. Memory accumulates what matters about users and successful workflows. Tool search adapts to the capabilities people actually use. Agent optimizer in Foundry Agent Service then closes the loop by analyzing agent behavior and generating improved instructions, skills, tool descriptions, and model configurations.

That is the larger goal of context engineering: not simply reducing prompt size or retrieval costs, but creating agents that improve with use. When the knowledge they draw from, the tools they discover, the procedures they follow, and the memories they retain all become better over time, an agent can become both more capable and more efficient without starting over.

Get started

If you’re building agents today, start by examining what enters the context window on every turn. Look at the documents being retrieved, the tools being exposed, the instructions being repeated, and the conversation history being carried forward. In many cases, improving those inputs has a larger impact on cost and quality than changing models.

Ground an agent in enterprise data with Foundry IQ: Connect a Foundry IQ knowledge base to an agent.

Give an agent one endpoint for its tools, and turn on tool search: Enable tool search in a toolbox.

Add memory so an agent carries context across sessions: Create and use memory in Foundry Agent Service.

Start building in Microsoft Foundry.

Microsoft Foundry

The enterprise AI platform to build, ground, and govern AI apps and agents at scale

Explore capabilities

Did you miss these posts in The Economics of Agent Optimization series?

AI cost management: From AI pilots to measurable ROI

AI cost optimization: How to lower AI spend

1 Command Line, Tool search: Finding the right tool at the right time, July 29, 2026.
The post The Economics of Agent Optimization: Context engineering for enterprise AI agents appeared first on Microsoft Azure Blog.
Quelle: Azure

Inside Microsoft’s marketing team: Scaling expertise with AI

Business leaders are facing a familiar challenge at an unfamiliar scale.

Every organization is being asked to move faster as markets change quickly, customer expectations continue to rise, and technology advances at a pace that can feel overwhelming. Teams are expected to deliver greater results, often with the same resources they had before.

AI is helping organizations meet those expectations. Some of the strongest examples I’ve seen revolve around scaling the judgment, strategy, and success measures that strong performers already set for themselves and their teams. AI agents apply that expertise consistently across a growing volume of work, helping them deliver more without sacrificing quality.

We’ve seen it firsthand on my team. As innovation cycles have accelerated, product launches have increased from a quarterly cadence to weekly—and sometimes even daily—events. Our teams are now supporting a growing volume of launches, up to 150% year over year.

To relieve the pressure, we’ve looked for places where AI can help teams at Microsoft find the right information faster, reduce repetitive coordination, and bring more consistency to work that depends on shared context. To do that, we used Microsoft Foundry, Microsoft’s platform for building and managing enterprise AI applications, to create agents grounded in business knowledge and embedded in the flow of work, helping our teams operate at greater scale while staying focused on the work where their expertise matters most.

Start building with Microsoft Foundry

Why context matters

One lesson became clear very quickly: AI is only as good as the data it has access to. General-purpose AI can generate content, but enterprise decisions depend on information spread across documents, workflows, business systems, communications, and institutional knowledge.

For us, Microsoft IQ helped connect that business context to our AI capabilities. Rather than asking employees to assemble information from multiple sources, agents could draw from the same knowledge people rely on every day to surface relevant information and support better decisions.

But IQ does more than ground AI in the right data. It helps connect the knowledge and workflows that shape how the business actually operates.

That shift changed the role AI could play. Instead of simply helping people find information, it could help teams work from a shared understanding of what’s happening across the business.

Learn how Microsoft IQ helps bring together people, data, knowledge, and workflows

Context alone wasn’t enough. The breakthrough wasn’t a single agent. It was creating a way for teams to build on what was already working.

As people shared successful agents and AI skills, expertise started becoming easier to reuse and scale. Ideas that began with one team could quickly create value for many others.

Microsoft Foundry became important because it allowed us to ground agents in organizational knowledge, connect them to existing workflows, and operationalize them beyond a single team.

In many ways, this reflects a broader lesson across AI adoption. As Jay Parikh recently wrote, “AI alone doesn’t transform a business. The system around it does.” The following examples show what that looked like inside our marketing organization:

const currentTheme =
localStorage.getItem(‘blogInABoxCurrentTheme’) ||
(window.matchMedia(‘(prefers-color-scheme: dark)’).matches ? ‘dark’ : ‘light’);

// Modify player theme based on localStorage value.
let options = {“autoplay”:false,”hideControls”:null,”language”:”en-us”,”loop”:false,”partnerName”:”cloud-blogs”,”poster”:”https://cdn-dynmedia-1.microsoft.com/is/image/microsoftcorp/1104726-MarThriveCustomerZero_tbmnl_en-us?wid=1280″,”title”:””,”sources”:[{“src”:”https://cdn-dynmedia-1.microsoft.com/is/content/microsoftcorp/1104726-MarThriveCustomerZero-0x1080-6439k”,”type”:”video/mp4″,”quality”:”HQ”},{“src”:”https://cdn-dynmedia-1.microsoft.com/is/content/microsoftcorp/1104726-MarThriveCustomerZero-0x720-3266k”,”type”:”video/mp4″,”quality”:”HD”},{“src”:”https://cdn-dynmedia-1.microsoft.com/is/content/microsoftcorp/1104726-MarThriveCustomerZero-0x540-2160k”,”type”:”video/mp4″,”quality”:”SD”},{“src”:”https://cdn-dynmedia-1.microsoft.com/is/content/microsoftcorp/1104726-MarThriveCustomerZero-0x360-958k”,”type”:”video/mp4″,”quality”:”LO”}],”ccFiles”:[{“url”:”https://azure.microsoft.com/en-us/blog/wp-json/bloginabox/v1/get-captions?url=https%3A%2F%2Fwww.microsoft.com%2Fcontent%2Fdam%2Fmicrosoft%2Fbade%2Fvideos%2Fproducts-and-services%2Fen-us%2Fazure%2F1104726-marthrivecustomerzero%2F1104726-MarThriveCustomerZero_cc_en-us.ttml”,”locale”:”en-us”,”ccType”:”TTML”}],”downloadableFiles”:[{“url”:”https://cdn-dynmedia-1.microsoft.com/is/content/microsoftcorp/1104726-MarThriveCustomerZero_transcript_en-us”,”locale”:”en-us”,”mediaType”:”transcript”},{“url”:”https://cdn-dynmedia-1.microsoft.com/is/content/microsoftcorp/1104726-MarThriveCustomerZero_audio_en-us”,”locale”:”en-us”,”mediaType”:”audio”}]};

if (currentTheme) {
options.playButtonTheme = currentTheme;
}

document.addEventListener(‘DOMContentLoaded’, () => {
ump(“ump-6a9834893027f”, options);
});

Raising the quality bar at scale

As our Microsoft Foundry business grew, so did the volume of content we needed to create. Our team now reviews and publishes more than 200 blog posts each year, maintaining a consistent quality bar increasingly dependent on a small number of subject matter experts. Much of their time was spent applying the same review criteria over and over again.

Rather than reviewing every draft from scratch, one of our content leaders documented the rubric she uses to evaluate a strong blog and refined it until it reflected the standards our team expected.

Using Microsoft Foundry, we translated that expert-defined rubric into a repeatable workflow that could identify gaps and opportunities before content reached a human reviewer. The capability was integrated directly into the content creation process, bringing instant feedback to every drafted post and making expert-defined standards available to every content creator.

Review cycles that once required substantial manual effort can now be completed in minutes, resulting in higher satisfaction and over 2,000 estimated hours saved annually across the team. More importantly, the approach demonstrates a broader pattern organizations can apply in many domains: use AI to apply established criteria at scale so experts can focus their time where judgment, coaching, and experience create the most value.

What we automated is consistency, not judgment. Our team set the bar based on our expertise; the AI agent reviews every post against that bar.

Validating messaging before it reaches customers

As the pace of innovation accelerated, one question kept coming up: would our messaging resonate with the customers we were trying to reach?

At Microsoft, we aim to keep the customer at the center of everything we do. That led us to look for ways to evaluate messaging before it reached customers, using more than internal opinions alone.

We applied that approach through AI Messaging Assistant (AMA), which helps evaluate messaging and positioning against different audience perspectives before going to market. Instead of relying only on internal opinions, teams can pressure-test whether a message is clear, relevant, and actionable for the stakeholders they are trying to reach.

Using Microsoft Foundry, we grounded AMA with a virtual congress of personas based on real customer conversations and extended it with the expertise, product knowledge, messaging guidance, and business context our teams rely on every day. That made it possible to move from a one-off AI experiment to a repeatable workflow where teams could evaluate messaging against a shared understanding of audience needs rather than rebuilding that understanding for every review.

The broader pattern is using AI to pressure-test important decisions before they reach customers, partners, or employees.

Microsoft scales customer intelligence with AMA to speed decisions, deliver ROI

Keeping teams aligned as the pace accelerates

As launch activity accelerated across our business, keeping teams aligned became harder than creating the work itself. New announcements arrived daily. Priorities shifted quickly. Information was spread across planning backlogs, documentation, meetings, and operational systems. Our marketers were spending too much time assembling context and not enough time acting on it.

In response, we started by writing down how our marketing work actually gets done, turning an unwritten process into a clear specification. With that in hand, we could sort the work: which parts required a marketer’s judgment, which could be automated, and which could be delegated to AI.

Using agents built on Microsoft Foundry, we connected the systems our teams already rely on, including planning backlogs, documentation, meeting signals, and other operational sources. Rather than manually gathering updates from dozens of places, teams can work from a real-time view of key developments, upcoming launches, and changes that affect go-to-market plans.

This transformed alignment from a manual effort into a repeatable workflow. Instead of spending time assembling information, teams spend more time understanding what changed, why it matters, and what actions to take next.

The challenge was never a lack of expertise. It was coordinating that expertise across a rapidly changing environment.

The outcome is not merely faster communication. It is better organizational alignment. When teams operate from the same base, decisions happen faster, handoffs become smoother, and organizations can respond more quickly to change.

Scaling expertise, reducing friction

Across each of these examples, the goal wasn’t automation for its own sake. The goal was making expertise available wherever it could create value. Looking back, the lesson wasn’t that a single AI capability changed how we worked. It was that building the right system around those capabilities allowed expertise, context, and judgment to scale across the team.

Technology will continue to evolve. The pace of business will continue to accelerate. But the differentiator remains the same: People provide the judgment. People set the strategy. People define success. AI helps them scale it.

Microsoft Foundry

Foundry helps teams create agents grounded in enterprise knowledge, connected to the tools people use every day, and designed for production workflows.

Ready to build with Foundry?

The post Inside Microsoft’s marketing team: Scaling expertise with AI appeared first on Microsoft Azure Blog.
Quelle: Azure

Introducing Azure Multicloud Interconnect for AWS

As organizations accelerate AI adoption and modernize their digital estates, applications, data, and infrastructure increasingly span multiple cloud environments. Customers are choosing the best platform for each workload, leveraging unique capabilities across providers to drive innovation, resilience, and business agility.

Yet while multicloud strategies have become commonplace, networking between cloud environments remains complex. Establishing private connectivity often requires customers to manually stitch together services, coordinate provisioning across providers, manage multiple operational processes, and navigate fragmented support experiences. What should be a straightforward connectivity decision can take weeks or months to implement and manage.

Explore Azure Multicloud Interconnect

Today, Microsoft and Amazon Web Services (AWS) are taking an important step toward simplifying that experience.

Microsoft Azure and AWS are excited to collaborate on a multicloud networking solution that uses both AWS Interconnect – multicloud and Azure Multicloud Interconnect for network interoperability, enabling customers to establish a private, high-performance private connectivity between Microsoft Azure and AWS through a streamlined, cloud-native experience.

This collaboration is built using the Open API specifications for network interoperability. Azure Multicloud Interconnect helps remove much of the complexity traditionally associated with multicloud networking. Customers can provision connectivity through an integrated experience while benefiting from enterprise-grade performance, resiliency, security, and operational simplicity.

Simplifying multicloud connectivity

Until now, organizations connecting Azure and AWS environments required careful planning, physical connectivity, routing configuration, provisioning coordination, monitoring, and lifecycle management to assemble and maintain multiple components across providers.

Azure Multicloud Interconnect fundamentally changes this model. With Microsoft and AWS collaborating using the standardized Open API specification, customers can establish dedicated private connectivity through a simplified experience that abstracts the underlying complexity of multicloud networking. Rather than focusing on infrastructure management, organizations can focus on delivering applications, moving data, and accelerating business outcomes.

The result is a high-bandwidth, more predictable path to deploying multicloud architectures for mission-critical workloads.

Designed for the AI era

Training and inference workloads frequently require access to data distributed across environments. Enterprises are increasingly architecting applications that span cloud boundaries while maintaining performance, security, and compliance requirements.

Azure Multicloud Interconnect is designed to support these evolving requirements with high-capacity private connectivity that extends to Azure Private Link, providing an end-to-end private path between the clouds.

This combination of high-performance connectivity and operational simplicity enables customers to move faster as they build the next generation of AI-enabled applications and services.

As AI transforms every industry, customers need the freedom to place data, applications, and infrastructure wherever it delivers the greatest business value. Azure Multicloud Interconnect helps make that possible by providing resilient, high-performance, private connectivity between Azure and AWS through a simplified, cloud-native experience. Together, we are reducing the complexity of multicloud networking and giving customers the scale, reliability, and agility they need to power the next generation of AI and data-driven innovation.

Customers told us they wanted a better way to connect workloads spanning AWS and Azure, and the old ways of doing it were clunky. With AWS Interconnect-multicloud and Azure Multicloud Interconnect, we’re proving what’s possible when both sides commit to a high bar: MACsec security out of the box, four-nines availability, and scalability at the click of a button.
—Robert Kennedy, VP of Network Services at AWS

Advancing an open multicloud future

Azure Multicloud Interconnect represents more than a new connectivity offering—it is a step toward a more open and interconnected cloud ecosystem.

Looking ahead, we see the opportunity to extend this model beyond a single cloud-to-cloud relationship. The same open API specification can help enable broader interoperability across hyperscale cloud providers, creating a more consistent experience for customers operating in increasingly diverse multicloud environments.

Customers can deploy connectivity at speeds up to 100 Gbps from day one at general availability, helping meet the needs of high-bandwidth applications while providing the foundation for future growth. As demand increases, capacity can expand dynamically, enabling organizations to scale without disrupting operations or redesigning their network architecture.

Beyond hyperscalers, this approach has the potential to simplify connectivity with network service providers and telecommunications carriers. By adopting a common interoperability model, cloud providers and carriers can work together to streamline last-mile connectivity, accelerate provisioning, and reduce operational complexity across the end-to-end customer journey.

Our long-term vision is an open ecosystem where hyperscalers, network service providers, and telecommunications carriers use a common interoperability framework to establish and operate connectivity through standardized APIs. Customers should be able to provision trusted, high-performance connectivity between clouds, metro networks, and enterprise locations with the same simplicity and automation they expect from modern cloud services.

To learn more about how to implement multicloud networking on Azure, please visit the in-depth blog or visit the Microsoft Learn page to get started. To read the AWS announcement, please visit their blog.

Get started with Azure Multicloud Interconnect

Explore technical guidance for implementing private, high-performance connectivity between Azure and AWS.

Read the in-depth blog

The post Introducing Azure Multicloud Interconnect for AWS appeared first on Microsoft Azure Blog.
Quelle: Azure

Managed PostgreSQL vs. self-hosted PostgreSQL: Key benefits and trade-offs

Summary
This post is for technical decision makers evaluating where to run production PostgreSQL workloads. It compares two valid operating models—self-managed PostgreSQL and a managed database service—through business and operational outcomes: control, engineering capacity, resilience, security, cost predictability, risk tolerance, and access to specialist expertise. The right choice depends on which responsibilities the organization needs to retain and which it is prepared to transfer to a service provider.

Why organizations choose PostgreSQL

PostgreSQL is a powerful, versatile open-source relational database used for workloads ranging from small applications to enterprise systems. It is a natural choice for teams that need a standards-based database with a broad ecosystem and extensive extensibility.

Managed vs. self-hosted PostgreSQL: Understanding the trade-offs

Once an organization chooses PostgreSQL, it must decide where and how to operate it: on its own infrastructure, in a virtual machine, or through a managed cloud service. Each model offers a different balance of control, responsibility, cost, and operational effort. Self-hosting provides direct control over the operating system (OS), PostgreSQL installation, and supporting infrastructure, but it also leaves the organization responsible for operating the complete platform. Managed PostgreSQL services, like Azure Database for PostgreSQL and Azure HorizonDB, can reduce selected infrastructure and platform responsibilities and help teams devote more engineering capacity to applications, data, and performance.

Discover migration services in Azure Database for PostgreSQL

Operational challenges emerge at scale

Self-hosting creates an ongoing ‘operational tax’: the time, expertise, and resources required to provision, secure, monitor, maintain, and recover the database platform. These activities are essential, but they do not usually differentiate the application or service the organization is building.

In a self-hosted scenario (e.g., running Postgres on a virtual machine (VM) or an on-premises server), the company’s engineering team is responsible for the entire stack:

Full-stack lifecycle management: Provisioning hardware, ensuring adequate power and cooling, maintaining datacenter infrastructure, installing the OS, and correctly configuring PostgreSQL.

Security hardening: Manually managing firewalls, OS-level security patches, and encryption at rest and in transit.

High Availability (HA): Setting up complex replication and failover mechanisms (like Patroni or Pacemaker), which are notoriously difficult to test and maintain.

Disaster recovery: Designing, automating, monitoring, and testing backups and recovery procedures, including point-in-time recovery. Achieving dependable recovery objectives requires sustained engineering effort and operational discipline.

Identity management: Manually managing database users and passwords, creating silos of credentials that are difficult to audit.

Operational tax consumes time that could be spent improving applications, delivering features, refining data models, or tuning the database’s performance.

When a managed PostgreSQL service may make sense

Managed database services change the operating model by transferring defined infrastructure and platform responsibilities to the cloud provider. This reduces the undifferentiated work required to keep the database platform available, secure, patched, and recoverable.

A managed PostgreSQL service is designed to reduce the operational tax associated with infrastructure and platform administration. The provider typically manages the underlying infrastructure, operating-system maintenance, service patching, and physical datacenter security. Customers remain responsible for their data, database configuration, access policies, application design, and workload performance.

Try Azure Database for PostgreSQL

When self-managed PostgreSQL may be the right fit

Self-management can be the preferred operating model when an organization requires operating-system access, specialized infrastructure, unsupported extensions, customized deployment patterns, or direct control over patching and change schedules. It may also fit organizations with mature PostgreSQL platform engineering, reliable automation, tested recovery practices, and sufficient on-call capacity. In these cases, the additional responsibility is an intentional trade-off for control and flexibility rather than avoidable burden.

Managed services and the shared responsibility model

A useful way to evaluate managed PostgreSQL services is through the lens of the cloud shared responsibility model. As organizations move from self-hosted environments to infrastructure-as-a-service (IaaS) and platform-as-a-service (PaaS) offerings, an increasing portion of the infrastructure and platform stack is operated by the cloud provider. In an IaaS deployment, organizations still manage virtual machines, operating systems, database software, and many operational processes. In a PaaS deployment, responsibility for operating systems and much of the underlying platform is transferred to the provider, reducing the operational effort required to keep the service secure, available, and up to date.

This shift does not eliminate responsibility. Organizations continue to own their data, identities, configurations, access management, compliance requirements, and application behavior regardless of the deployment model. However, by reducing responsibility for operating-system management, physical infrastructure, platform maintenance, and portions of the security stack, managed PostgreSQL services can significantly reduce the operational tax associated with running database platforms at scale.

For example, Azure Database for PostgreSQL is a PostgreSQL PaaS offering that transfers responsibility for infrastructure management, operating-system maintenance, patching orchestration, and platform operations to Microsoft while allowing organizations to retain control over their data, database configuration, access policies, and workload design. This balance enables teams to focus more effort on application delivery, data architecture, and performance optimization rather than routine platform administration. This follows the same shared-responsibility principles that apply across cloud PaaS services.

How managed PostgreSQL services can help

Managed PostgreSQL services replace selected manual infrastructure and platform tasks with managed capabilities and configurable automation. The comparisons below show how responsibilities associated with self-hosting can be simplified or transferred to the service, reducing the Operational Tax on engineering teams.

1. Automated lifecycle management

Self-hosted: Teams must monitor for security alerts, manually download patches, and plan downtime for both the OS and the database.

Managed service: The provider typically manages operating-system maintenance and service updates, including minor PostgreSQL version updates. Customers can often configure preferred maintenance windows for planned maintenance, while major-version upgrades may remain customer-initiated so compatibility can be validated and the schedule controlled.

2. High availability and resilience by design

Self-hosted: Manual HA requires managing multiple VMs, replication lag, and witness nodes. Failover is rarely 100% reliable without constant testing.

Managed service: High availability is commonly available as a service configuration, with the provider maintaining standby capacity and orchestrating failover under the service commitment. Read replicas are a separate capability for scaling read-intensive workloads and should be evaluated independently from the high-availability design.

3. Intelligent storage and recovery

Self-hosted: Exhausted disk capacity can cause a high-severity incident. Teams must provision and monitor storage, maintain backup automation, validate backup targets, and plan capacity increases or disk replacement.

Managed service: Storage growth, automated backups, and point-in-time restore are commonly built into the platform. Retention periods, storage limits, redundancy options, and scaling behaviour vary by provider, region, service tier, and configuration.

4. Enterprise security and identity

Identity and access management comparison

FeatureSelf-hostedManaged serviceAuthenticationManual password rotation and managing separate, siloed credential stores.Native Microsoft Entra ID integration for centralized identity management.Security postureHigher risk of credential leaks due to manual handling and static passwords.Passwordless authentication support, significantly reducing the attack surface.AdministrationDBAs must manually sync organizational users with database roles.Use Microsoft Entra identities and groups to centralize access management and align database access with organizational identity processes. Database permissions still require administration.

Evaluation checklist

Use these questions to evaluate the two operating models:

Do we have the infrastructure and PostgreSQL expertise to operate the platform reliably?

Which responsibilities must remain under direct organizational control?

What availability, recovery, security, and compliance outcomes are required?

How much operational variation and incident risk can the organization absorb?

Can the platform scale to support new workloads?

Which work should database specialists own, and which responsibilities can be transferred to a managed service?

How important are predictable operating costs, standardized controls, and faster deployment?

The choice: Strategic allocation of resources

Choosing between self-managed PostgreSQL and a managed service is a strategic allocation of control, engineering capacity, and operational risk. Self-management can be appropriate when direct infrastructure access, specialized capabilities, or deep customization outweigh the cost of owning the complete platform. A managed service can be appropriate when standardized operations, resilience, security integration, and faster delivery outweigh the need for low-level control.

Technical decision makers should select the model that aligns with business priorities, risk tolerance, regulatory obligations, and the organization’s ability to operate PostgreSQL reliably at the required scale. For teams that choose a managed model on Azure, Azure Database for PostgreSQL provides a fully managed PostgreSQL service with configurable high availability, maintenance, backup, scaling, and security capabilities. Azure HorizonDB provides a cloud-native PostgreSQL option for mission-critical workloads that benefit from independently scalable compute and storage, rapid read scale-out, and built-in zone resilience.

Ready to evaluate the next step?

First validate required extensions, operating-system access, availability and recovery objectives, security controls, regional availability, performance needs, and internal support capacity. If a managed service on Azure fits those requirements, the following resources can help you evaluate Azure Database for PostgreSQL and plan a migration.

Recommended starting point: Migration Service in Azure Database for PostgreSQL flexible server – Azure Database for PostgreSQL

Hands-on training: Migrate to Azure Database for PostgreSQL flexible server – Training

Cloud shared responsibility model: Shared responsibility in the cloud

Online migration tutorial: Migrate Online, from an Azure VM or an On-Premises PostgreSQL to Azure Database for PostgreSQL flexible server, Using the Migration Service in Azure – Azure Database for PostgreSQL

Offline migration tutorial: Migrate Offline, from an Azure VM or an On-Premises PostgreSQL to Azure Database for PostgreSQL flexible server, Using the Migration Service in Azure – Azure Database for PostgreSQL

Simplify Your PostgreSQL Migration

Learn how the migration service in Azure Database for PostgreSQL helps move databases to flexible server with minimal complexity and downtime.

Start Your Migration

The post Managed PostgreSQL vs. self-hosted PostgreSQL: Key benefits and trade-offs appeared first on Microsoft Azure Blog.
Quelle: Azure

The Economics of Agent Optimization: Four ways to lower the cost

This blog post is the second of a four-part series called The Economics of Agent Optimization which shares the strategies, capabilities, and proof points to help you optimize agent costs and run AI as a managed investment system on Microsoft Foundry. The first post set out the three decisions that system rests on: optimize each request at runtime, optimize each workflow over time, and govern spend continuously. This post takes the first, the one that touches every dollar you will ever spend on AI.

An agent is a loop around a model. It plans, calls a tool, reads the result, and reasons again, so a single completed outcome can take a dozen model requests. That is why the number the business cares about is the cost of a successful outcome, not the price of a token.

Every turn in that loop is still one model request, and each request carries decisions about the model, the offer it runs on, what gets reused, and what the model is told. When those decisions are right, the saving repeats on every turn. That is why agent optimization starts here.

Learn how Microsoft Foundry can help you save time and money

The most expensive habit in production AI

Most AI applications are built the same way. In the prototype, you pick the strongest model available, put everything the model might need into the prompt, and confirm the idea works. That is the correct instinct for a prototype; the problem is what happens next. The prototype’s defaults quietly become the production architecture, and a pattern designed to answer “can this work?” becomes responsible for answering “can this scale economically?”

Two things break at that point. First, AI workloads are not uniform. A single application mixes intent classification, extraction, formatting, summarization, and genuine multi-step reasoning—AI workloads vary enormously in complexity. Routing all of them to one frontier model means overpaying on the majority of requests that never needed that capability.

Second, one outcome is many requests. A prototype pays for a single call. An agent pays for the whole loop, so anything wasteful gets multiplied. That is true of tokens, and it is more true of mistakes. An agent that takes a wrong turn calls the wrong tool and loops to recover, burning tokens on turns that should never have happened and still landing on a weaker answer. Cost per outcome is set as much by the turns you avoid as by the tokens in each one.

In production, the goal is not to minimize tokens. It’s to reduce the cost of a successful outcome while maintaining quality, safety, and latency. Every runtime decision must balance those factors together, which is why the economics of a request come down to four decisions.

Four levers you control at runtime

Microsoft Foundry gives you four levers for making those tradeoffs deliberately, rather than accepting the ones your prototype happened to choose. Each can be adopted on its own, measured against your quality bar, and reversed if the tradeoff does not hold.

LeverFoundry capabilityModels and offersModel router, deployment types, provisioned throughput, batch, fine-tuning.CachingPrompt caching, semantic caching through the AI Gateway in Azure API Management.Prompt and agent optimizationPrompt optimizer, agent optimizer across instructions, skills, tool descriptions, and model selection.Observability and evaluationFoundry observability and evaluation, agent traces, Azure budgets, alerts, and cost tagging.

1. Send each request to the right model

The principle is simple: optimize the outcome based on the task complexity. Routine requests should not pay frontier-model economics, while complex requests should not sacrifice quality simply to save tokens.

Model router in Foundry Models removes that tradeoff. It assesses each incoming request and dispatches it to the most suitable underlying model in real time, behind a single endpoint and a single deployment. Routing modes let you prioritize cost, quality, or a balance of the two. Model subsets, which now align with Azure Policy, constrain routing to an approved allow-list where a compliance boundary applies. Built-in failover moves a request to the next best model when one is unavailable, so routing also buys resilience.

const currentTheme =
localStorage.getItem(‘blogInABoxCurrentTheme’) ||
(window.matchMedia(‘(prefers-color-scheme: dark)’).matches ? ‘dark’ : ‘light’);

// Modify player theme based on localStorage value.
let options = {“autoplay”:false,”hideControls”:null,”language”:”en-us”,”loop”:false,”partnerName”:”cloud-blogs”,”poster”:”https://cdn-dynmedia-1.microsoft.com/is/image/microsoftcorp/1119598-ModelRouterInFoundryModels_tbmnl_en-us?wid=1280″,”title”:””,”sources”:[{“src”:”https://cdn-dynmedia-1.microsoft.com/is/content/microsoftcorp/1119598-ModelRouterInFoundryModels-0x1080-6439k”,”type”:”video/mp4″,”quality”:”HQ”},{“src”:”https://cdn-dynmedia-1.microsoft.com/is/content/microsoftcorp/1119598-ModelRouterInFoundryModels-0x720-3266k”,”type”:”video/mp4″,”quality”:”HD”},{“src”:”https://cdn-dynmedia-1.microsoft.com/is/content/microsoftcorp/1119598-ModelRouterInFoundryModels-0x540-2160k”,”type”:”video/mp4″,”quality”:”SD”},{“src”:”https://cdn-dynmedia-1.microsoft.com/is/content/microsoftcorp/1119598-ModelRouterInFoundryModels-0x360-958k”,”type”:”video/mp4″,”quality”:”LO”}],”ccFiles”:[{“url”:”https://azure.microsoft.com/en-us/blog/wp-json/bloginabox/v1/get-captions?url=https%3A%2F%2Fwww.microsoft.com%2Fcontent%2Fdam%2Fmicrosoft%2Fbade%2Fvideos%2Fproducts-and-services%2Fen-us%2Fazure%2F1119598-modelrouterinfoundrymodels%2F1119598-ModelRouterInFoundryModels_cc_en-us.ttml”,”locale”:”en-us”,”ccType”:”TTML”}],”downloadableFiles”:[{“url”:”https://cdn-dynmedia-1.microsoft.com/is/content/microsoftcorp/1119598-ModelRouterInFoundryModels_transcript_en-us”,”locale”:”en-us”,”mediaType”:”transcript”},{“url”:”https://cdn-dynmedia-1.microsoft.com/is/content/microsoftcorp/1119598-ModelRouterInFoundryModels_audio_en-us”,”locale”:”en-us”,”mediaType”:”audio”}]};

if (currentTheme) {
options.playButtonTheme = currentTheme;
}

document.addEventListener(‘DOMContentLoaded’, () => {
ump(“ump-6a8f3cdb126fd”, options);
});

The same request can carry very different economics depending on how it is deployed, and this is the lever teams most often leave untouched. Organizations must decide:

Where data is processed (Global, Data Zone, or Regional).

How throughput is purchased (pay-per-token, provisioned capacity).

Which workloads truly require interactive responses.

Foundry provides multiple deployment options that allow these choices to align with business requirements. Most workloads can start with standard deployments, which provide the greatest flexibility and cost-efficient pay-as-you-go pricing. Interactive applications that require faster and more consistent response times can benefit from priority processing, while high-volume workloads with predictable demand can achieve better economics through Provisioned Throughput Units (PTUs), with overflow traffic handled through pay-as-you-go capacity. Large asynchronous workloads such as document processing, classification, and evaluation runs are often best suited for Batch deployments, which provide up to 50% lower costs for work that doesn’t require immediate responses.

Even within a single application, different experiences often benefit from different deployment strategies. Developer-facing tools that can tolerate some latency variability may run efficiently on Standard deployments. Interactive chat experiences may warrant priority processing, while agentic applications with sustained throughput demands can maximize value with PTUs. Background tasks such as document analysis, knowledge extraction, and large-scale classification can move to Batch without affecting the end-user experience, reducing cost simply by selecting the deployment model that matches the workload.

Fine-tuning is the advanced version of this lever. Where routing picks among existing models, fine-tuning changes what a smaller model can do, teaching it your task, tone, or format well enough to match a larger model on that job. The payoff is a lower rate and shorter prompts. Reach for it when behavior is stable and volume is high enough to earn back the effort.

2. Stop paying for the same tokens twice

Agents are highly cache effective. The same system instructions, tool schemas, and policy text are re-sent on every turn, so an agent that takes 10 turns pays for that prefix 10 times. Prompt caching lets a previously processed prefix be reused rather than reprocessed. Cache reads are billed at a discount to normal input pricing on standard deployments, and can be discounted up to 100% on provisioned deployments. Latency improves alongside cost.

Getting value from it is mostly a matter of prompt architecture, and the rule is straightforward: stable content first, volatile content last. Put system instructions, tool definitions, and few-shot examples at the top, and user input, retrieved chunks, and turn history at the bottom. Caching depends on an exact match at the start of the prompt, so anything that changes per request, such as a timestamp or a user’s name, has to sit below that block. Put it at the top and the cache never matches.

Caching works above the prompt too. When you deploy a gateway in front of the Foundry inference APIs, it’s important to choose a semantic-cache-aware gateway such as the AI Gateway in Azure API Management. It can maintain session affinity to the same endpoints, helping maximize cache effectiveness while matching near-duplicate requests across sessions and users. Deterministic tool results can be cached in your own store with a time-to-live tuned to how often the data changes.

3. Optimize the prompt, then optimize the agent

If model choice sets the rate, the instruction sets the volume. It is also the cheapest thing to fix, because it ships without touching infrastructure. The practices that cut tokens are the same ones that improve answers: lead with the task rather than burying it after a wall of context, be specific about the output you want and how long it should be, and use a few well-chosen examples in place of paragraphs of explanation. Then keep what accumulates across turns under control:

Summarize completed conversations instead of replaying full transcripts.

Scope tool definitions to only the tools relevant to the task.

Store working state in external memory and retrieve it only when needed.

Foundry now automates the hand-tuning this used to take. Prompt optimizer rewrites an agent’s system instructions using prompt-engineering best practices and shows its reasoning for each change, so you can steer it, run it again, and apply the result in a click.

Agent optimizer in Foundry Agent Service goes further and closes the loop. It runs your agent against a dataset of real tasks, generates candidate configurations, scores each one, and ranks them so you can promote the winner. It can change instructions, skills, tool descriptions, and model selection, and the dataset can come from your own agent traces.

4. Make it visible with observability and evaluation

You cannot tune what you cannot see, and you cannot claim a saving you did not measure. Observability in Foundry supplies the per-request signals that make the other three levers safe to pull: input and output tokens, cache hit rate, latency, the model that actually served the request, and the evaluation scores that say whether quality held.

Two numbers matter here. Cost per request tells you whether the cheaper path still cleared the bar. Cost per completed outcome tells you what the business actually paid, across every turn and retry it took to get there. An optimization that lowers the first while raising the number of turns has made things worse, and only the second will show it.

Evaluation turns that visibility into permission to change things. Measure cost, latency, and task success together, and keep a standing evaluation set that every optimization has to clear before it ships. Those same traces and evaluation sets are what agent optimizer consumes, so the work pays twice. Pair them with budgets, alerts, and cost tagging in Azure so a regression arrives as a notification rather than a surprise at month end.

How this adds up to a hill climb

None of these levers is a one-time saving. Together they form a loop that gets cheaper and better every time you go around it, which is what Microsoft AI means by building a hill-climbing machine: improving continuously, cycle after cycle, through better data and sharper evaluation.

Model and offer decides where each request runs, and fine-tuning turns a proven task into a permanently cheaper one.

Caching lowers the cost of every cycle, which is what lets you run the loop often enough to matter.

Prompt and agent optimization generates the next candidate and proves it against your evaluation set.

Observability and evaluation tells you where you are and whether the last change held.

The climb is a cycle with no fixed start, though most teams enter it at measurement. Traces become evaluation datasets. Those datasets drive the optimizer. Optimizer results show which tasks are stable enough to fine-tune. Fine-tuned models change what the router should choose, and the new routing produces fresh traces.

Get started

Read the first post in this series, on moving from AI pilots to measurable ROI: The Economics of Agent Optimization: from pilots to measurable returns .

Deploy model router and compare it against your current baseline.

Watch the token economics episode on Microsoft Mechanics for a hands-on look at these levers in action.

Microsoft Foundry

Learn more about runtime optimization

Start building with Foundry

Did you miss these posts in The Economics of Agent Optimization series?

AI cost management: From AI pilots to measurable ROI

The post The Economics of Agent Optimization: Four ways to lower the cost appeared first on Microsoft Azure Blog.
Quelle: Azure

The patch window is collapsing: Why security needs a new control plane

For decades, cybersecurity defenders have relied on a relatively straightforward model: a vulnerability is disclosed, security teams assess exposure, test available fixes, deploy patches into production, and ultimately close the risk before attackers can exploit it at scale. 

That model increasingly reflects a world that no longer exists.

Today’s enterprises operate thousands of interconnected workloads across hybrid and multicloud environments. Mission-critical applications power revenue-generating services, customer experiences, and core business operations that cannot simply be taken offline whenever a security update becomes available. At the same time, vulnerabilities are becoming more visible, more widely distributed, and more rapidly weaponized than ever before.

The result is a growing gap between how quickly organizations can safely remediate vulnerabilities and how quickly adversaries can exploit them. It is time to rethink how the industry approaches security during the critical period between disclosure and remediation. 

The patch window has collapsed 

Traditional vulnerability management was built on the assumption that defenders could move faster than attackers. In many cases, they could.

When a vulnerability was disclosed, organizations had time to understand the issue, assess affected systems, test patches, coordinate change windows, and deploy fixes before widespread exploitation occurred.

Today that timeline is rapidly shrinking.

Modern attack campaigns operate at internet scale. Security research, public disclosures, proof-of-concept exploits, and threat intelligence circulate globally within hours. A vulnerability announced in the morning can become the focus of active scanning and exploitation efforts by the afternoon.

Meanwhile, the operational realities of enterprise environments have not changed. Organizations still must: 

Understand the vulnerability and its business impact. 

Identify affected systems across large estates. 

Evaluate dependencies and compatibility concerns. 

Validate fixes in test environments. 

Coordinate deployment schedules. 

Monitor for regressions and operational risk. 

These are not signs of inefficiency. They are necessary safeguards for business-critical environments. The challenge is that while defensive processes continue to require days or weeks, offensive timelines are increasingly measured in hours. 

That creates one of the most dangerous periods in modern cybersecurity: the window between awareness and remediation.

AI is expanding the defender’s challenge 

AI is helping organizations modernize operations, accelerate development, and improve security outcomes. But the same technological advances are also changing the economics of offensive operations.

Historically, transforming a newly disclosed vulnerability into an effective attack often required extensive manual research and deep technical expertise. Security researchers and attackers alike needed to analyze documentation, understand exploit conditions, study affected software, and develop attack techniques.

Many of those steps can now be accelerated.

AI-assisted workflows can help analyze vulnerability disclosures, identify likely attack paths, evaluate technical dependencies, and summarize complex technical information far more quickly than traditional manual processes.

As these capabilities become more accessible, the timeline between disclosure and exploitation continues to compress. The result is a structural imbalance. 

Defenders remain responsible for protecting entire environments that may include thousands of servers, applications, databases, containers, and network assets. Attackers only need to identify a single viable path to exploitation. 

This asymmetry is driving organizations to ask an increasingly important question: What happens before the patch is deployed?

Why existing security approaches fall short 

The security industry has invested heavily in improving visibility.

Organizations today have access to more vulnerability data, threat intelligence, analytics, and detection capabilities than ever before. Security platforms can rapidly identify affected systems, prioritize remediation, and alert defenders to emerging threats.

These capabilities are essential. But awareness alone does not reduce exposure. Many organizations find themselves in a position where they know exactly which systems are vulnerable but cannot immediately patch them. 

For example, a business-critical application may require extensive validation before updates can be deployed. A manufacturing system may depend on software that cannot be taken offline during production hours. A regulated environment may require additional testing and approval processes before changes can be implemented.

In these situations, the challenge is not identifying risk. The challenge is reducing risk while remediation is still underway.

Visibility, detection, and prioritization help organizations understand the problem. They do not necessarily provide a mechanism for containing that risk immediately.

As attack timelines continue to compress, the industry needs a complementary approach focused on exposure reduction rather than simply exposure awareness.

Why the network is emerging as the fastest control plane 

When a workload cannot immediately defend itself, another layer must help provide protection. Increasingly, organizations are looking to the network. 

Unlike endpoint-based controls, network-level protections operate around workloads rather than inside them. This distinction becomes particularly important during periods of elevated risk.

The network already understands communication patterns, connectivity requirements, trust relationships, and traffic flows. It sits at a strategic position where organizations can influence how systems interact with one another without necessarily modifying the applications themselves.

This creates opportunities to reduce exploitability while remediation efforts are underway. Network-enforced protections can help: 

Restrict access to vulnerable systems. 

Limit exposure to potential attack paths. 

Reduce opportunities for lateral movement. 

Segment high-risk assets. 

Contain potential blast radius. 

Adjust controls dynamically as new information becomes available. 

Perhaps most importantly, network controls can often be implemented significantly faster than enterprise software patches can be validated and deployed.

The objective is not to avoid patching. The objective is to create a meaningful layer of defense during the period when patching has not yet been completed.

As AI compresses the time between vulnerability disclosure and exploitation, organizations need a defensive layer that can act immediately, without waiting for every workload to be patched, every application to be modified, or every endpoint agent to understand a new threat.

The network is uniquely positioned to become that control point: it already sits in the path of communication, has visibility across heterogeneous workloads, and can enforce protections consistently across large cloud estates without changing the applications themselves. More importantly, network controls can increasingly move beyond simple IP, port, and signature-based blocking toward context-aware, adaptive enforcement that constrains the specific behavior an exploit depends on while preserving legitimate traffic.

Consider an HTTP/2 denial-of-service vulnerability: the safest interim guidance may be to disable HTTP/2 entirely until systems are patched, but that can carry significant application and performance impact. A more precise network and workload-aware response could instead bound the exploitable behavior—limiting concurrent streams, tightening request constraints, or rate-limiting abusive connection patterns—while keeping the service available. This is why the network is becoming more than a connectivity layer: it can serve as a programmable, ubiquitous enforcement fabric that buys organizations the most valuable commodity during a zero-day—the time to patch safely.

In an era where vulnerabilities may be weaponized within hours, every day of risk reduction matters.

The rise of adaptive security 

The next evolution of cybersecurity is unlikely to rely solely on static policies or manual response processes. Modern environments are simply too large, dynamic, and interconnected. 

Organizations increasingly need security systems capable of understanding risk, evaluating context, and adapting protections as conditions change. This shift points toward a broader industry trend: adaptive security. 

Adaptive security systems aim to move beyond predefined rules toward continuously improving risk management. Rather than treating every vulnerability equally, they seek to understand the specific conditions that make a flaw exploitable and determine the most effective way to reduce exposure. At a high level, these systems must solve three critical challenges. 

First, they must understand the vulnerability itself. 

This requires ingesting information from security advisories, vulnerability disclosures, threat intelligence, exploit research, and other sources to develop a meaningful understanding of how a threat operates.

Second, they must correlate that understanding with real-world environments. 

A vulnerability only becomes a material risk when specific systems, configurations, connectivity paths, and exposure conditions exist. Understanding this context is essential to determining actual risk.

Third, they must translate intelligence into action. 

Insight without enforcement provides limited value. The ultimate goal is to reduce exposure through controls that can be applied quickly, consistently, and at scale.

AI is expected to play a significant role throughout this process, not merely as an analytical tool, but as an enabling technology that helps security systems understand complex relationships and make informed decisions faster than would otherwise be possible.

Looking at the future of cybersecurity

The cybersecurity industry has spent decades improving vulnerability management, patch deployment, and security operations. Those investments remain essential and will continue to be foundational elements of every organization’s security strategy. But the environment around us is changing.

Attackers are moving faster. Infrastructure is becoming more complex. AI is compressing timelines across the entire threat landscape. In this new reality, organizations cannot rely on patching alone. 

The future of cybersecurity will depend on an organization’s ability to reduce risk during the time between disclosure and remediation. Success will come from combining strong patch management practices with compensating controls capable of responding at machine speed.

The organizations that thrive will be those that treat security as a continuous, adaptive process rather than a sequence of point-in-time responses. The fundamental question is no longer whether vulnerabilities will emerge. They will. 

The question is how effectively organizations can protect themselves while they work to eliminate them.

As the patch window continues to collapse, the industry will need new approaches that complement traditional remediation strategies, reduce exposure quickly, and help defenders regain the one resource that has become increasingly scarce in modern cybersecurity: time.

Microsoft is investing in new and innovative capabilities able to provide immediate protection from the storm, buying organizations the time they need to safely validate and deploy a permanent patch without exposing their environment to unnecessary risk.
The post The patch window is collapsing: Why security needs a new control plane appeared first on Microsoft Azure Blog.
Quelle: Azure