AnnouncementBuilt for unpredictable AI demand: Atlas Infinite and MongoDB 9.0 are here. Read more >  >>
AnnouncementMeet the intelligent data platform built for the AI era. Read more > >>
NewNow in public preview: Atlas Agent Engine, the secure way to run AI agents at scale. Read more > >>
Blog home
arrow-left

Context Engineering Has a Retrieval Issue: Introducing rerank-3 and Native Reranking

September 30, 2026 ・ 6 min read

Every team building AI agents discovers one fundamental truth: without the right information, an agent cannot reliably answer a question or take the correct action.

As large language model context windows grew larger, it was tempting to assume that just giving an agent more information would improve its accuracy. In practice, more context makes agents slower and more expensive, while also reducing accuracy by burying useful signals among irrelevant information.

That’s why the success of an agent now relies on context engineering, the practice of selecting, organizing, and presenting an agent with only the information it needs to complete a task effectively.

And the core of that work happens in retrieval. The question isn't how much you can fit in context. It's which ten things out of ten million you choose.

MongoDB’s intelligent data platform brings together Voyage AI’s best-in-class embedding and reranking models with a platform for building, operating, and evolving complex retrieval systems. With today’s introduction of the Rerank 3 series and $rerank in the aggregation pipeline, developers can improve the quality and ordering of retrieved context in the same system where their operational data lives. The result is a simpler way to build retrieval systems that give agents the context they need to take action.

The retrieval stack

Effective retrieval requires a coordinated system to find useful information, determine what matters most, and efficiently deliver it to the agent.

The strategy determines how the system searches its data. Keyword search can surface exact terms, while vector search finds conceptually similar content. Filters and hybrid approaches add precision and control, ensuring that only the relevant context is retrieved.

Models determine what relevance looks like for your data. A stronger embedding model helps vector search identify better candidates by understanding conceptual similarity and query intent. A reranker allows first-stage retrieval to optimize for recall, then separates the strongest matches from the candidate set.

The right engine brings these disparate technologies together, keeping data current and queries efficient as the system runs in production.

Without a unified platform to manage retrieval, the work gets complicated fast. You might use one provider for embeddings, another for reranking, and multiple database vendors for your operational, search, and vector data. That means more integrations, sync jobs, and API calls to maintain. It can also create silent problems: your vector search index may use one embedding model version while your application uses another, hurting relevance without causing an explicit error.

Figure 1. The disconnected retrieval pipeline.

A diagram illustrating a disconnected retrieval pipeline, showing arrows flowing from a query and unstructured data through separate embedding models, vector search, and rerankers before reaching an LLM for a grounded response.

MongoDB brings these pieces together in a single platform. It supports the search strategies developers need, including text, vector, hybrid, and filtered retrieval, and pairs them with Voyage AI’s state-of-the-art embedding models and rerankers. Automated Embedding keeps vectors synchronized as data changes, while native rerank lets developers refine results in the same aggregation pipeline. The result is a single query that can search operational data, apply the right retrieval strategy, rerank the results, and return the context an agent needs.

Figure 2. The unified retrieval pipeline powered by MongoDB.

A diagram showing the unified retrieval pipeline powered by MongoDB and Voyage AI, where a query flows through an integrated system containing embedding models, unstructured data, vector search, and rerankers directly to an LLM for a grounded response.

The models

Today’s introduction of rerank-3 and rerank-3-lite offers improved accuracy across all tasks, with the largest gains on long documents and code. rerank-3 outperforms Cohere Rerank v4.0 Pro by 2.72% and Qwen3-Reranker-8B by 3.02% overall. If you're on rerank-2.5 today, upgrading is a one-line model-name change. Same API, same price.

rerank-3 improves relevance no matter how the initial candidate set is retrieved. The table below shows the accuracy lift across several common first-stage retrieval methods.

Table 1. Accuracy comparisons.

First stageNDCG@10 lift from $rerank
BM25 (lexical only)+20.56%
OpenAI v3 large+12.77%
voyage-4-large+5.08%

The engine

MongoDB is the operational retrieval engine underneath all of this, and every strategy runs natively against the same data your application already writes to. Lexical search catches identifiers, error codes, and exact product names. Vector search catches concepts and paraphrases. Hybrid retrieval combines them, and metadata filters keep the candidate set small enough to stay fast. None of it requires moving data into a separate system or holding two copies in sync.

Automated Embedding handles the part that usually breaks. Vectors are generated and refreshed at the field level in near real time, with token consumption and index health visible in the Atlas UI. There's no embedding job to own, and no chance for a split between the model version your index was built on and the model version your queries run against.

$rerank, available today, adds the final ordering step to that same pipeline. It's an index-less stage that can take whatever candidate set the previous stage produced and rescore it with rerank-3 or rerank-3-lite. No separate vendor, no external API call, no reassembly step in application code — which matters when an agent is making several retrieval calls per task rather than one.

Figure 3. $rerank in action.

A diagram illustrating $rerank in action, showing how initial search results for "corkboard," "flipchart," and "dry-erase surface" are rescores and reordered by relevance scores of 0.9, 0.8, and 0.7 respectively.

Gravity: What it looks like when this works

Gravity is building an ad network that optimizes advertising for the AI era. Relevance isn't a nice-to-have there; it's the product. Their problem wasn't recall. Earlier embedding and reranking models returned relevance scores clustered too tightly together, so Gravity couldn't confidently separate a strong match from a weak one.

“The biggest thing we found in our benchmarks was the ability of Voyage AI to be discriminatory to candidates that are not a good match and really promote the ones that are a good match,” said Leo Martinez, Gravity's cofounder and CTO. Switching to Voyage AI cut retrieval latency by roughly 50% and improved accuracy as request volume climbed.

Closing the gap

Retrieval quality is one layer of what MongoDB is building for agents to run on. MongoDB 9.0 and Atlas Infinite are the performance and elastic-scale layer, absorbing the bursty traffic that agentic workloads create. Atlas Agent Engine is the memory and governance layer, giving agents persistent context and guardrails across sessions; retrieval is powered by Atlas Vector Search and Voyage AI. An agent decides what to retrieve before it decides what to do. Faster infrastructure over bad retrieval just means acting on the wrong thing sooner.

Getting retrieval right no longer requires complex architecture. Automated Embedding handles the embeddings, $rerank handles the ordering, and both run against your operational data directly in MongoDB. The models are also available standalone through the Atlas Embedding and Reranking API.

megaphone
Next Steps

Get started with Automated Embedding and Native Reranking here. For everything we announced at MongoDB.local NYC—as well as the latest product updates—visit our What's New page.  

MongoDB Resources
MongoDB Products|Atlas Learning Hub|MongoDB University|Documentation|MongoDB Events