Skip to content
Not available in this workspace
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Benchmarks
  • Apps
  • Discover
  • Models
  • Collections
  • Providers
  • Pricing
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Support
  • Works With OR
  • Data

Developer

  • Documentation
  • API Reference
  • Developer Platform
  • Status

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube
Collections/Rerank Models

Best Rerank Models for Search and RAG

Model rankings updated August 2026 based on real usage data.

Rerank models reorder candidate documents, passages, or search results by relevance, sharpening the retrieval step in semantic search, retrieval-augmented generation (RAG), and recommendation pipelines. This collection ranks reranking models by their usage on OpenRouter over the past week. The current top models are Llama Nemotron Rerank VL 1B V2 (free), rerank-2.5-lite, and rerank-2.5. Compare pricing, latency, and quality to find the best fit for your search or RAG pipeline.

Browse All ModelsCompare Models

Top Rerank Models on OpenRouter

Favicon for nvidia

NVIDIA: Llama Nemotron Rerank VL 1B V2 (free)

18.8B tokens

Llama Nemotron Rerank VL 1B V2 is a 1.7B multimodal reranking model from NVIDIA. It evaluates the relevance of document images and text against user queries, designed for vision RAG pipelines handling charts, tables, infographics, and mixed-media documents. Functions as a cross-encoder that accepts text queries paired with image, text, or combined document inputs, delivering approximately 6-7% recall improvements over embedding-only baselines on visual document retrieval benchmarks.

by nvidia10K context$0/M input tokens$0/M output tokens
Favicon for voyageai

VoyageAI by MongoDB: rerank-2.5-lite

10.4B tokens

rerank-2.5-lite is a reranker optimized for both latency and quality, delivering a 7.16% improvement in retrieval accuracy over Cohere Rerank v3.5 across 93 datasets. It also outperformed Cohere Rerank v3.5 by 10.36% on the Massive Instructed Retrieval Benchmark (MAIR). The model supports a combined context length of 32K tokens per query–document pair, including up to 8K tokens for the query, enabling more accurate retrieval over longer documents. Additionally, rerank-2.5-lite supports instruction following, allowing users to guide relevance scoring through natural language prompts. Learn more about rerank-2.5-lite here: blog.voyageai.com/2025/08/11/rerank-2-5

by voyageai32K context$0.02/M tokens
Favicon for voyageai

VoyageAI by MongoDB: rerank-2.5

2.92B tokens

rerank-2.5 is a cutting-edge reranker optimized for quality, delivering a 7.94% improvement in retrieval accuracy over Cohere Rerank v3.5 across 93 datasets. It also outperformed Cohere Rerank v3.5 by 12.70% on the Massive Instructed Retrieval Benchmark (MAIR). The model supports a combined context length of 32K tokens per query–document pair, including up to 8K tokens for the query, enabling more accurate retrieval over longer documents. Additionally, rerank-2.5 supports instruction following, allowing users to guide relevance scoring through natural language prompts. Learn more about rerank-2.5 here: https://blog.voyageai.com/2025/08/11/rerank-2-5

by voyageai32K context$0.05/M tokens
Favicon for qwen

Qwen3 Reranker 8B

468M tokens

Qwen3 Reranker 8B is a text reranking model from Alibaba Cloud built on the Qwen3 architecture. It evaluates query-document pairs to produce relevance scores for use in retrieval and RAG pipelines. Supports 100+ languages and programming languages, with instruction-aware reranking that allows customizing scoring criteria per task. Offers strong performance on multilingual benchmarks including MTEB, CMTEB, and MMTEB.

by qwen41K context$0.20/M tokens
Favicon for cohere

Cohere: Rerank 4 Pro

Cohere's AI search foundation model for enhancing the relevance of information surfaced within search and RAG systems. Features a 32K context window, multilingual support across 100+ languages, no data pre-processing required, and state of the art performance with low latency.

by cohere33K context$0.0025/search
Favicon for cohere

Cohere: Rerank 4 Fast

Cohere's AI search foundation model for enhancing the relevance of information surfaced within search and RAG systems. Features a 32K context window, multilingual support across 100+ languages, no data pre-processing required, and high performance with lowest latency.

by cohere33K context$0.002/search
Favicon for cohere

Cohere: Rerank v3.5

Rerank v3.5 is designed to reorder search results for improved relevance. It supports multi-aspect and semi-structured data reranking over 100+ languages. Ideal for refining results from semantic or keyword search pipelines.

by cohere4K context$0.001/search

Explore more collections

  • Free Models
  • Discounted Models
  • Coding
  • Roleplay
  • Vision Models
  • Tool Calling
  • OpenClaw
  • Image Models
  • Video Models
  • Audio Models
  • Text-to-Speech
  • Speech-to-Text
  • Embedding Models
  • Distillable Models
  • All collections