Use cases
Choosing an AI API for RAG (retrieval-augmented generation)
Retrieval-augmented generation sends retrieved passages plus a question, so requests carry a lot of input. The model must stick to the passages and admit when they don't answer the question.
What RAG (retrieval-augmented generation) needs
- Low input price for large retrieved context
- Faithfulness to the supplied passages
- Citations back to sources
Models that fit
| Model | Maker | List price (in / out per 1M) | Through Somnus | Why |
|---|---|---|---|---|
| Gemini 3.8 Flash | $0.75 / $3.75 | $0.38 / $1.88 | cheap large inputs | |
| Claude Haiku 4.5 | Anthropic | $1 / $5 | $0.5 / $2.5 | follows the sources closely |
| Claude Sonnet 5 | Anthropic | $2 / $10 | $1 / $5 | hard multi-document questions |
Somnus column uses the minimum 2× credit; larger top-ups go further.
Cost example
1,000 requests of about 1,500 input and 400 output tokens on Gemini 3.8 Flash cost about $2.63 at list price — roughly $1.31 with Somnus credit.
Tips
- Retrieve fewer, better passages rather than many weak ones.
- Number the passages and ask the model to cite them.
- Tell the model to answer 'not in the sources' when appropriate.
Common mistakes
- Sending the top 50 chunks when the top 5 would do.
- Not instructing the model to refuse when context is missing.
Example
from openai import OpenAI
import os
client = OpenAI(
api_key=os.environ["SOMNUS_API_KEY"],
base_url="https://gateway-production-c837.up.railway.app/v1",
)
resp = client.chat.completions.create(
model="claude-sonnet-5",
messages=[
{"role": "system", "content": "You are an expert assistant. Be precise and concise."},
{"role": "user", "content": "Using only the passages below, when was the company founded?"}
],
)
print(resp.choices[0].message.content)Questions
What is the cheapest model for RAG (retrieval-augmented generation)?
Of the models suggested here, Gemini 3.8 Flash has the lowest list price ($0.75 / $3.75 per 1M tokens). Test it on your own examples before committing.
Can I switch models later?
Yes. With Somnus you change the model name; the rest of the code stays the same.
Run this code with free credit
Create an account and get $5 of free API credit — enough to try every example on this page.
Get $5 free creditRelated
Browse more: Alternatives · Compare · Use cases · Developers · Solutions · Tools