Developers

Build summarization in Python

Summarising is input-heavy: you send a long document and get a short answer back. Input price per million tokens drives the bill. Below is a working Python starting point using the official openai Python package.

Python example

python
from openai import OpenAI
import os

client = OpenAI(
    api_key=os.environ["SOMNUS_API_KEY"],
    base_url="https://gateway-production-c837.up.railway.app/v1",
)

resp = client.chat.completions.create(
    model="claude-sonnet-5",
    messages=[
      {"role": "user", "content": "Summarise this meeting transcript in five bullet points."}
    ],
)
print(resp.choices[0].message.content)

Models that fit

ModelMakerList price (in / out per 1M)Through SomnusWhy
Gemini 3.8 FlashGoogle$0.75 / $3.75$0.38 / $1.88low input price
DeepSeek V4 FlashDeepSeek$0.22 / $0.66$0.11 / $0.33lowest input price in the table
Claude Sonnet 5Anthropic$2 / $10$1 / $5nuanced or high-stakes documents

Somnus column uses the minimum 2× credit; larger top-ups go further.

Making it production-ready in Python

  • Read the key from an environment variable, never hard-code it.
  • Use the async client (AsyncOpenAI) for web servers.
  • Split very long documents and summarise the summaries.
  • Tell the model the target length and audience.

Avoid

  • Choosing a model on output price when your cost is almost all input.
  • Not asking the model to quote or cite, which hides errors.

Questions

Do I need a special SDK for summarization in Python?

No — the official openai Python package is enough, because Somnus uses the OpenAI chat-completions format.

Run this code with free credit

Create an account and get $5 of free API credit — enough to try every example on this page.

Get $5 free credit

Related

Browse more: Alternatives · Compare · Use cases · Developers · Solutions · Tools

WindowsStart with $5 free