
Engineering 8 min read· 15 Aug 2026· By Engineering
Reranking for better RAG — cheaper than a bigger LLM
Retrieve 50, rerank to top 5, feed to the model. The 20% effort that gets 80% of the gain.
Embedding search recalls the right document 70% of the time. A cross-encoder reranker pushes that to 90%+ for pennies.
Rule of thumb: retrieve 50, rerank to 5, prompt on those 5. Better than retrieving 20 directly, every time.