NVIDIA Nemotron 3.5 Lightning model is now available on Amazon SageMaker JumpStart

NVIDIA’s Nemotron 3.5 Lightning is now available on Amazon SageMaker JumpStart, giving AWS customers access to the fastest open model in its class for persistent agent workloads and rapid task execution.
Nemotron 3.5 Lightning is engineered for persistent agents and high-throughput enterprise automation across domains including personal assistants, financial document processing, cybersecurity triage, and telecom operations. Built on a hybrid Mixture-of-Experts (MoE) architecture with 30B total parameters and just 3B active per forward pass, it achieves up to 4x the throughput (~410 tokens/sec) and 30% faster task completion over comparable models. Distilled from Nemotron 3 Ultra, it handles up to 1M tokens of context via DFlash speculative decoding and integrates directly with popular agent harnesses. The model is fully open-trained on open datasets thereby allowing enterprises to post-train for their own tools, workflows, and policies, and deploy with complete ownership across edge, on-premises, or cloud infrastructure.
With SageMaker JumpStart, customers can deploy this model in a few clicks to power their specific AI workloads.
To get started with these models, navigate to the SageMaker JumpStart model catalog in the SageMaker console or use the SageMaker Python SDK to deploy the models to your AWS account. For more information about deploying and using foundation models in SageMaker JumpStart, see the Amazon SageMaker JumpStart documentation.
Quelle: aws.amazon.com

LocateAnything-3B, Qwen-AgentWorld-35B-A3B, and Qwen3.5-122B-A10B models now available on Amazon SageMaker JumpStart

NVIDIA’s LocateAnything-3B, Qwen’s Qwen-AgentWorld-35B-A3B, and Qwen’s Qwen3.5-122B-A10B models are now available on Amazon SageMaker JumpStart, expanding the portfolio of foundation models available to AWS customers. These three models bring specialized capabilities spanning visual grounding, agent environment simulation, and large-scale multimodal reasoning, enabling customers to deploy high-performance, scalable AI solutions on AWS infrastructure.
These models address different enterprise AI challenges with specialized capabilities:
LocateAnything-3B is optimized for fast, high-quality visual grounding and object localization from natural language instructions. It uses a Parallel Box Decoding (PBD) framework that decodes bounding boxes and points as atomic units in a single step, preserving geometric coherence and unlocking substantial parallelism. It enables precise object localization, dense detection, and point-based localization across diverse domains in both Enterprise Intelligence and Physical AI applications.
Qwen-AgentWorld-35B-A3B excels in simulating agent environments across seven interaction domains: tool calling, search, terminal, software engineering, Android, web, and OS interaction. It is the first language world model to cover all seven domains within a single model, predicting next environment states given an agent’s action and interaction history via long chain-of-thought reasoning—trained on over 10 million real-world interaction trajectories.
Qwen3.5-122B-A10B provides high-performance multimodal reasoning with production-friendly efficiency. It features 122B total parameters with only 10B activated per token through a hybrid architecture integrating Gated Delta Networks with sparse Mixture-of-Experts (256 experts), delivering strong reasoning, coding, agents, and visual understanding performance with a native 262K context window and minimal latency overhead.
With SageMaker JumpStart, customers can deploy any of these models with just a few clicks to address their specific AI use cases.
To get started with these models, navigate to the SageMaker JumpStart model catalog in the SageMaker console or use the SageMaker Python SDK to deploy the models to your AWS account. For more information about deploying and using foundation models in SageMaker JumpStart, see the Amazon SageMaker JumpStart documentation.
 
Quelle: aws.amazon.com

Amazon GameLift Streams Now Offers Service-managed Shader Caching

Amazon GameLift Streams now manages shader cache capture and distribution for your applications. You capture a shader cache from a stream session, and the service automatically makes it available for future sessions across your streaming locations. No application changes are required. Capturing shader caches can help reduce loading times and visual stuttering, during the session. With service-managed shader caching, you designate a stream session for capture and run your application to generate the cache. Amazon GameLift Streams then replicates the cache to compatible stream groups and locations, and loads it automatically in future sessions. You can monitor shader cache status and storage size using the ListApplicationShaderCaches API or the Amazon GameLift Streams console. The feature supports Linux (Ubuntu 22.04), Proton, and Windows Server 2022 runtimes. You are charged for storage of the latest version of each shader cache. For pricing details, visit the Amazon GameLift Streams pricing page. For supported Regions, see the AWS Region table.
Quelle: aws.amazon.com

Amazon EC2 High Memory U7i instances now available in AWS South America (São Paulo) region

Amazon EC2 High Memory U7in-24TB instances (u7in-24tb.224xlarge) are now available in AWS South America (São Paulo) region. U7i instances are part of the AWS 7th generation and are powered by custom fourth-generation Intel Xeon Scalable processors (Sapphire Rapids). U7in-24TB instances offer 24 TiB of DDR5 memory, enabling customers to scale transaction processing throughput in a fast-growing data environment.
U7in-24TB instances deliver 896 vCPUs and support up to 100 Gbps of Amazon EBS bandwidth for faster data loading and backups, 200 Gbps of network bandwidth, and ENA Express. U7i instances are ideal for customers running mission-critical in-memory databases like SAP HANA, Oracle, and SQL Server.
To learn more about U7i instances, visit the High Memory instances page.
Quelle: aws.amazon.com

GLM-5.2 FP8, NVIDIA-Nemotron-Nano-12B-v2 and GLM-OCR models now available on Amazon SageMaker JumpStart

Z.ai’s GLM-5.2 FP8, NVIDIA’s Nemotron-Nano-12B-v2, and Z.ai’s GLM-OCR models are now available on Amazon SageMaker JumpStart, expanding the portfolio of foundation models available to AWS customers. These three models bring specialized capabilities spanning long-horizon agentic engineering, efficient hybrid reasoning, and advanced document understanding, enabling customers to deploy high-performance, scalable AI solutions on AWS infrastructure.
GLM-5.2 FP8 is optimized for long-horizon tasks and agentic engineering workflows such as full-cycle software development from requirements to deployment. It delivers a substantial leap in long-horizon task capability over its predecessor GLM-5.1 and, for the first time, provides a truly usable 1M-token context window, enabling it to handle project-level engineering context, execute long-running tasks reliably, follow engineering standards consistently, and complete full development workflows in a single task.
NVIDIA-Nemotron-Nano-12B-v2 excels in unified reasoning and non-reasoning tasks with high inference throughput, making it ideal for enterprise applications requiring both accuracy and efficiency. It uses a hybrid Mamba-2 and Transformer architecture with a 128K context length, generating reasoning traces before concluding with final responses. Its compact 12B parameter design achieves comparable or better accuracy than leading open models while delivering up to 6x higher inference throughput.
GLM-OCR provides accurate, fast, and comprehensive document understanding for complex real-world materials including scanned PDFs, handwritten notes, dense academic papers with formulas, multi-column tables, code documentation, and multilingual text. This 0.9B-parameter multimodal model reconstructs structure, tables, and formulas into clean Markdown, JSON, or LaTeX, with latency low enough for real-time services and edge devices—ideal for large-scale document processing and invoice extraction workflows.
With SageMaker JumpStart, customers can deploy any of these models with just a few clicks to address their specific AI use cases.
To get started with these models, navigate to the SageMaker JumpStart model catalog in the SageMaker console or use the SageMaker Python SDK to deploy the models to your AWS account. For more information about deploying and using foundation models in SageMaker JumpStart, see the Amazon SageMaker JumpStart documentation.
Quelle: aws.amazon.com