Beyond the benchmark: How an adaptive approach drives scientific discovery

For research and development (R&D) organizations, the promise of agentic AI is not a better one-time answer. It is a new way to explore complex scientific and engineering problems: pursuing multiple hypotheses, validating them against evidence, learning from what does not work, and adapting their approach as new information becomes available.

This unique nature of the agentic discovery process has been a core area of research for Microsoft, and a design principle for Microsoft Discovery, our platform for organizations embracing Frontier R&D.

Learn how Microsoft Discovery explores science

Measuring adaptive AI for scientific discovery

A new benchmark result shows how that opportunity is becoming real. On Agent’s Last Exam, a demanding evaluation of long-running, tool-using professional tasks, Microsoft Discovery Engine with CLIO (Cognitive Loop via In-Situ Optimization) achieved higher scores than the other agentic harnesses evaluated across three scientific domains: 61.6% in health and medicine, 75.2% in physical sciences, and 64.6% in life sciences.

This result builds on Microsoft’s core research into what makes agentic discovery distinctive. CLIO enables independent reasoning paths to explore a problem, compare and share learning, and resolve the strongest trajectory into a single evidence-backed result. The system can determine when to keep exploring, change strategy, use a different model, or bring a domain expert into the loop.

The CLIO benchmark blog post describes this adaptive reasoning approach in depth. More broadly, this core innovation for scientific discovery, powered by agentic AI, is available to R&D organizations in every industry and the scientific community with Microsoft Discovery not only as a research breakthrough, but as a foundation for real R&D work.

Explore Microsoft Discovery

Why scientific discovery requires adaptive reasoning

Many of the hardest scientific and engineering challenges do not have a clearly defined workflow or a known answer. A researcher may need to navigate incomplete evidence, competing objectives, specialized tools, and changing constraints. A materials team may be balancing performance, safety, cost, and manufacturability. A life sciences team may need to connect literature, proprietary data, models, and experimental evidence before deciding what to validate next. An engineering team may need to search a vast design space without sacrificing physical fidelity or traceability.

In these settings, a single model response is not enough. Practitioners need systems that can reason over time, preserve evidence, challenge assumptions, and work within the tools, data, governance, and review processes their experts already use. Just as importantly, they need to understand how a conclusion was reached and where human judgment should enter the process.

Microsoft Discovery was designed as an enterprise platform for agentic R&D, combining the scientific mindset of hypothesis, experimentation, and refinement with the engineering rigor of problem decomposition, structured execution, and reproducibility. CLIO strengthens that foundation with a more adaptive reasoning loop and a diverse model ecosystem, while allowing researchers to use a diverse model ecosystem and multiple reasoning paths.

Try Microsoft Discovery app

From benchmarks to real-world impact

The greater opportunity extends beyond benchmark rankings into real research environments. Discovery Engine with CLIO has already supported work that discovered a novel organic redox flow battery. The same approach has potential across design simulation (like for silicon chips), formulation and process optimization (for example in manufacturing and CPG), materials and molecular discovery (which can drive sustainability and drug discovery), and lab automation, areas where organizations need to shorten research cycles, without sacrificing rigor or traceability.

Agentic discovery does not replace scientists and engineers. It expands what they can explore, helps them learn faster from evidence, and gives them a more systematic and transparent way to move from an idea toward an outcome that experts can evaluate and validate.

Realizing the enormous opportunity to redefine R&D requires a platform built for the tools, data, governance, and review processes researchers already use. Microsoft Discovery was designed with that need in mind: to bring agentic discovery to researchers and scientists in R&D organizations across every industry and throughout the scientific community.

We are still early in this journey, but this benchmark milestone demonstrates what becomes possible when AI is built for the way discovery actually happens: iteratively, collaboratively, and adaptively. I look forward to seeing what organizations, researchers, and partners discover next.

Adaptive AI for scientific discovery

Learn how Microsoft Discovery uses adaptive, agentic approaches to explore complex scientific and engineering challenges, helping R&D teams accelerate innovation and uncover new possibilities.

Read the blog

The post Beyond the benchmark: How an adaptive approach drives scientific discovery appeared first on Microsoft Azure Blog.
Quelle: Azure

AWS Systems Manager now diagnoses more issues that cause EC2 instances to be unmanaged

Today, AWS Systems Manager extends its diagnosis capability to identify six additional categories of issues that can prevent Amazon EC2 instances and hybrid-activated nodes from becoming managed by Systems Manager. An instance must be managed by Systems Manager before you can patch it, run commands, connect with Session Manager, or collect inventory, and when an instance is unmanaged the cause can be difficult to isolate. The diagnosis previously covered network connectivity, and it now also identifies issues with IAM permissions, SSM Agent version, instance status checks, operating system configuration, Default Host Management Configuration, and hybrid activation.
With this broader coverage, more of your instances return a specific, actionable cause instead of an unidentified result, so you can bring your fleet under management faster. You run a diagnosis across your instances in the Systems Manager unified console experience, and Systems Manager reports the specific issues it finds in each category. Every diagnosed issue comes with step-by-step guidance to help you resolve it, and for some issues you can also run an AWS Systems Manager Automation runbook from the console to remediate the issue directly.
AWS Systems Manager diagnosis capability is available in all AWS Regions that are enabled by default. The capability runs as AWS Systems Manager Automation runbooks, so you pay standard Automation usage charges for the runbooks you run; see AWS Systems Manager pricing for details. To learn more, see the AWS Systems Manager User Guide, or visit the AWS Systems Manager product page.
Quelle: aws.amazon.com

Amazon Bedrock Managed Knowledge Base now supports Confluence Data Center as a native data source connector

AWS announces the Confluence Data Center data source connector for Amazon Bedrock Managed Knowledge Base, a fully managed retrieval-augmented generation (RAG) service. Customers running self-hosted Confluence Data Center instances can now crawl blogs and pages from their Confluence spaces directly into their managed knowledge base. Previously, bringing Confluence Data Center content into Bedrock Knowledge Bases required building custom ingestion pipelines—now, you provide your instance credentials, and the connector handles data crawling, metadata extraction, and incremental sync automatically. The Confluence Data Center connector gives your AI agents access to the institutional knowledge your teams already maintain in Confluence, including wiki pages and blog posts across spaces. You can use filters to scope crawls to specific spaces or content types, ensuring only relevant content is ingested and keeping your knowledge base focused and cost-efficient. This makes it easy to power internal assistants grounded in engineering documentation, runbooks, or team knowledge bases hosted on Confluence Data Center. To learn more, see Confluence Data Center data source connector in the Amazon Bedrock User Guide. For more information about Amazon Bedrock Managed Knowledge Base, visit the Amazon Bedrock Knowledge Bases product page.
Quelle: aws.amazon.com

Amazon Bedrock Managed Knowledge Base adds APIs and console support for debugging document-level access control

AWS announces the CheckIngestedDocumentAcl and GetIngestedDocumentAcl APIs for Amazon Bedrock Managed Knowledge Base, giving customers a self-service way to debug document access issues and audit document-level permissions. When a user doesn’t see an expected document in retrieval results for ACL-enabled data sources, it can be difficult to determine whether the cause is an access control misconfiguration or something else entirely. These new APIs and the accompanying console experience close that gap, enabling you to quickly diagnose and resolve permission issues without opening a support case.
CheckIngestedDocumentAcl lets you verify whether a specific user has access to a given ingested document, while GetIngestedDocumentAcl returns the full ACL attached to a document so you can audit exactly what permissions are configured and catch misconfigurations. The console also introduces a new Document Access Control section on the data source details page, where you can check access by entering a document ID and user email or retrieve a document’s full ACL by document ID. Together, these capabilities give administrators the visibility they need to manage access control at scale across enterprise knowledge bases.
To learn more, see CheckIngestedDocumentAcl and GetIngestedDocumentAcl in the Amazon Bedrock API Reference. For more information about Amazon Bedrock Managed Knowledge Base, visit the Amazon Bedrock Knowledge Bases product page.
Quelle: aws.amazon.com

AWS Lambda now supports Graviton5-powered EC2 instances on Lambda Managed Instances

AWS Lambda now supports AWS Graviton5-powered C9g, C9gd, M9g, and M9gd instances on Lambda Managed Instances. You can now run your Lambda functions on the latest generation of Graviton processors, delivering up to 25% better compute performance compared to Graviton4-powered instances.
AWS Lambda Managed Instances lets you run Lambda functions on your AWS EC2 instances while maintaining Lambda’s operational simplicity. With Lambda Managed Instances, you can access specialized compute configurations and drive cost efficiency through EC2 pricing advantages, without managing infrastructure. Lambda Managed Instances fully manages all infrastructure tasks, including instance lifecycle, OS and runtime patching, built-in routing, load balancing, and auto-scaling based on configurable parameters – so you can focus on writing code. Starting today, you can leverage Lambda Managed Instances with the latest Graviton5-powered instances, which deliver up to 25% better compute performance compared to the previous generation Graviton4-powered instances.
To get started, you can specify the desired Graviton5 instance type (C9g, C9gd, M9g, or M9gd) when you create a capacity provider. When you set the instance type to default, Lambda automatically includes Graviton5 instances in the list of instances it chooses from, based on your function’s configured architecture, memory size, and memory-to-vCPU ratio.
AWS Lambda Managed Instances supports C9g, C9gd, M9g, and M9gd instance types in all AWS Regions where both Lambda Managed Instances and these EC2 instances are available. To learn more, visit the Lambda Managed Instances documentation and AWS Lambda pricing. 
Quelle: aws.amazon.com

AWS Lambda now supports 90-minute function timeout on Lambda Managed Instances

AWS Lambda now supports a 90-minute function timeout for asynchronous and event source mapping (ESM) invocations on Lambda Managed Instances (LMI), a 6x increase from the previous 15-minute limit. You can now run data processing, media transcoding, financial calculations, AI inference, and batch workloads on Lambda for jobs that require longer continuous execution, without re-architecting your applications.
Customers use Lambda to build serverless applications like event-driven processors, API backends, and data processing pipelines. For data-intensive workloads like media transcoding, financial calculations (such as Monte Carlo simulations), and AI inference that need longer continuous execution, Lambda’s 15-minute function timeout limit required customers to adopt architectural workarounds. With today’s launch, you can configure a function timeout of up to 90 minutes for asynchronous and ESM invocations on Lambda Managed Instances. Lambda Managed Instances lets you process multiple concurrent requests per instance, access specialized compute configurations, and drive cost efficiency through EC2 pricing advantages, without managing infrastructure. The increased function timeout also applies to invocations within Lambda durable functions, which allow you to checkpoint and replay steps for longer-running invocations. When invoked asynchronously, a multi-step durable execution can run for up to 1 year.
You can configure up to a 90-minute function timeout for asynchronous and ESM invocations via the AWS Lambda Console, AWS CLI, Lambda APIs, Infrastructure as Code tooling, or the Agent Toolkit for AWS. Synchronous invocations retain the existing 15-minute maximum timeout. This feature is available in all AWS Regions where Lambda Managed Instances is available.
To learn more about configuring the 90-minute function timeout, see the Lambda developer guide. For combining extended timeouts with checkpoint-and-replay resilience, see the durable functions documentation. For pricing details, see AWS Lambda Pricing. To learn more about AWS Lambda, visit aws.amazon.com/lambda. 
Quelle: aws.amazon.com