← Docs

Quick start

One call before your LLM: fewer tokens in, fewer tokens out, same substance. Your LLM key never touches AgentReady.

1

Get a key

Get one in 30 seconds (free during beta, no card). Keys start with ak_. Store it as an environment variable:

export AGENTREADY_API_KEY=ak_...
2

Squeeze your messages

POST the messages you were about to send. Pick how hard to squeeze the input and whether the answer should be short too.

cURL
curl https://agentready.cloud/v1/compress \
  -H "Authorization: Bearer $AGENTREADY_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "messages": [{"role": "user", "content": "Hi! Could you please explain, in as much detail as possible, why my build fails?"}],
    "compression_level": "caveman",
    "output": "caveman"
  }'
Response
{
  "messages": [
    {"role": "system", "content": "Answer terse like caveman: no articles, filler, pleasantries, hedging. Fragments OK. Keep code, commands, numbers, names, URLs, errors exact; code, JSON and requested formats complete."},
    {"role": "user", "content": "Explain, in detail, why my build fails?"}
  ],
  "stats": {
    "original_tokens": 20, "compressed_tokens": 9, "tokens_saved": 11, "reduction_percent": 55.0,
    "compression_level": "caveman", "output_mode": "caveman", "output_rule_tokens": 46, "sent_tokens": 55
  },
  "output_rule": "Answer terse like caveman: ..."
}

original/compressed tokens count your content only; the output rule is reported separately in output_rule_tokens. On tiny prompts the rule can add more input tokens than it saves, the payoff is in the much shorter (and pricier) answer.

3

Call your LLM with what comes back

Python · OpenAI
import os, requests
from openai import OpenAI

def squeeze(messages):
    try:
        r = requests.post(
            "https://agentready.cloud/v1/compress",
            headers={"Authorization": f"Bearer {os.environ['AGENTREADY_API_KEY']}"},
            json={"messages": messages, "compression_level": "caveman", "output": "caveman"},
            timeout=3,
        )
        return r.json()["messages"] if r.ok else messages
    except requests.RequestException:
        return messages  # never block your app

reply = OpenAI().chat.completions.create(
    model="gpt-4o-mini",
    messages=squeeze([{"role": "user", "content": user_text}]),
)
Node · Anthropic
const r = await fetch('https://agentready.cloud/v1/compress', {
  method: 'POST',
  headers: { Authorization: `Bearer ${process.env.AGENTREADY_API_KEY}`, 'Content-Type': 'application/json' },
  body: JSON.stringify({
    provider: 'anthropic',
    system: 'You are the support agent for Acme.',
    messages: [{ role: 'user', content: userText }],
    compression_level: 'caveman',
    output: 'lite',
  }),
});
const { system, messages } = await r.json();
const reply = await anthropic.messages.create({ model: 'claude-sonnet-5-5', max_tokens: 1024, system, messages });

Request parameters

FieldTypeDefaultWhat it does
messagesarrayrequiredChat messages, OpenAI format. Only text content is touched; images, tool calls and short messages pass through.
compression_levelstring"standard"Input lever: light · standard · aggressive · caveman. caveman removes greetings, polite wrappers, hedges, filler and articles.
outputstring"off"Output lever: off · lite · caveman · ultra. Adds a short system rule so the model answers terse while keeping code, numbers and requested formats complete.
systemstring—Anthropic-style system prompt. It is squeezed too, and the output rule is appended to it. Returned as system.
providerstring—Set "anthropic" to always receive the output rule in system instead of a system message.

What is never changed

Fenced code blocks (byte-for-byte, indentation included), inline code, URLs, numbers and units, error messages and stack frames, double-quoted text and negations (not, never, no). The output rule tells the model the same: terse prose, complete code and data.

Choosing levels