AnnouncementBuilt for unpredictable AI demand: Atlas Infinite and MongoDB 9.0 are here. Read more >  >>
AnnouncementMeet the intelligent data platform built for the AI era. Read more > >>
NewNow in public preview: Atlas Agent Engine, the secure way to run AI agents at scale. Read more > >>
Blog home
arrow-left

Intelligent Workload Management: Keeping Your Clusters Available When It Matters

October 8, 2026 ・ 5 min read

Modern applications need to stay available even when traffic spikes suddenly or queries consume disproportionate resources. MongoDB Atlas is designed for high availability, but availability in practice is about more than surviving host failures or a replica set election. It’s also about how a cluster behaves when demand exceeds what it can safely process.

Last week, we announced Intelligent Workload Management (IWM) as GA in MongoDB 9.0. IWM is a managed availability capability that keeps Atlas clusters responsive and predictable even under overload. IWM queues transient spikes, keeps both short-running operations moving and long-running operations stable, and protects cluster health without requiring any application changes.

Overload events often unfold faster than infrastructure can scale. Even with auto-scaling enabled, provisioning additional capacity takes time. In the gap between a traffic spike and new capacity coming online, a cluster can become unstable if every request is treated the same way and clients retry aggressively. IWM is built to close that gap by preserving cluster health, optimizing performance, and turning what would be an unpredictable outage into a controlled, observable event.

“

At Coinbase, trading volume can spike in minutes, and our systems have to scale in real time, or customers feel it. Intelligent Workload Management gives us confidence that our clusters will stay responsive during peak trading windows. 

”
Krishnan Seshadrinathan, VP of Engineering, Coinbase

For sustained overload, customers can optionally enable Load Shedding to reject excess operations. IWM helps manage pressure while a spike passes or capacity is added; it does not add capacity or replace auto-scaling and capacity planning.

Figure 1. Intelligent Workload Management adapts to keep clusters responsive and predictable.

Infographic for Intelligent Workload Management (IWM) in MongoDB 9.0, featuring the headline "When Atlas hits its limit, it doesn't fall over, it adapts" and listing four key benefits: Atlas-managed overload protection, keeping clusters responsive and predictable, buying time for auto-scaling or operator action, and turning unpredictable outages into controlled, observable events.

Managing workload pressure while capacity catches up

Traffic spikes and resource-intensive operations can increase demand faster than infrastructure can scale. Even when auto-scaling is enabled, provisioning additional capacity takes time. Aggressive client retries can add pressure while a cluster is already busy.

IWM helps manage workload pressure during that period. It works with or without auto-scaling: when auto-scaling is enabled, IWM helps keep the cluster stable as capacity comes online; without it, IWM manages work at the cluster’s provisioned capacity. For sustained or growing demand, adding capacity may still be necessary.

How IWM responds to workload pressure

When a cluster is healthy, IWM is effectively invisible. When demand exceeds what the cluster can safely process, IWM manages pressure in stages:

  • It limits incoming work and queues surplus requests: By default, IWM queues surplus requests rather than proactively rejecting them. This can help absorb transient spikes, though queued requests may take longer to complete while resources are constrained. Queueing changes when work is handled, not the amount of work the application ultimately asks the cluster to do.
  • It helps keep short-running operations moving: A short-running operation, such as a typical read or write, and a long-running operation, such as a large reporting query, do not place the same demands on a cluster. When resources are constrained, IWM keeps short-running operations moving while managing long-running operations more conservatively.
  • It helps manage resource pressure: Memory fragmentation, cache imbalance, and storage exhaustion can all destabilize a healthy cluster. IWM adjusts cluster memory and cache management when it detects unhealthy patterns and can pause new writes while still permitting reads when storage usage reaches critical levels.

These protections help preserve cluster health during overload. They do not make capacity unlimited or replace good indexing, capacity planning, or auto-scaling.

Default protections and optional Load Shedding

IWM’s default protections are enabled at no additional cost on eligible dedicated Atlas clusters running MongoDB 9.0. They require no application changes or driver upgrades. By default, IWM manages overload without actively rejecting excess work.

Load Shedding is a separate, optional control for sustained overload. If enabled, it can reject excess operations, and those rejections are visible to your application as overload errors. Load Shedding is available at no additional charge for eligible M30+ dedicated replica-set Atlas clusters running MongoDB 9.0.

The default IWM protections work with any MongoDB driver. If you enable Load Shedding, use a supported backpressure-aware driver so your application can interpret overload signals and apply bounded retries with backoff. Check the supported driver matrix before enabling Load Shedding.

Figure 2. The difference between the default (queuing) and the Load Shedding setting.

Comparison of IWM modes showing Default Queuing on the left (lower-friction upgrade, graceful degradation, best starting point) versus Opt-In Load Shedding on the right (bounded latency and goodput, application-visible behavior, requires readiness).

Understand IWM interventions with Observability

In Atlas, review the Operation Rate Limiting and Connection Rate Limiting metrics for queued, rejected operations or connections. When Load Shedding is enabled, rejected operations appear under Operation Rate Limiting. Configure or review the IWM alert condition to track relevant IWM activity, and check the Activity Feed for IWM-related events. Together, these Atlas surfaces help you assess whether pressure is continuing and whether to review workload patterns or add capacity. Learn more about IWM observability in Atlas.

For teams evaluating optional Load Shedding, Atlas also provides a test to validate driver compatibility, application handling of overload errors, and operational visibility before relying on the control in production. When tested on active test or development clusters, Atlas randomly rejects a small percentage of operations regardless of load. Teams can follow rejected operations in the metrics chart and review the results after the test ends. Because applications may encounter overload errors during testing, teams should use these tests only to verify client behavior. Teams shouldn’t run tests on a production cluster. Learn how to test Load Shedding.

For details on IWM behavior and how to respond to overload, see the IWM documentation.

Getting started

  1. Confirm that your cluster is an eligible dedicated Atlas cluster running MongoDB 9.0.

  2. Decide whether auto-scaling should add capacity during sustained overload. IWM helps manage pressure while additional capacity comes online; it does not replace auto-scaling or capacity planning.

  3. If you enable Load Shedding, check the supported driver matrix and confirm that your application uses bounded retries with backoff.

  4. Review Atlas observability dashboards to understand when IWM intervenes and which workloads are affected.

megaphone
Next Steps

Ready to keep your clusters more predictable during overload events? Read the IWM documentation, explore Atlas auto-scaling, or sign up for MongoDB Atlas to start building today. 

MongoDB Resources
MongoDB Products|Atlas Learning Hub|MongoDB University|Documentation|MongoDB Events