Product

Route, trace, and grade LLM traffic.

Glimrel is a core AI/ML control plane for teams that already ship agents and RAG apps. The gateway keeps models up. The trace store keeps the story. The evaluator keeps the bar. Hosted cloud, or a license that stays on your floor. Not an agent platform.

Multi-model gateway

One base URL across Claude, GPT-4o, Gemini, Llama 3, Nova, and Groq. Fallbacks, caches, and budgets per key. Provider outages fail over without a code change.

Nested LLM traces

Spans for completions, tools, retrieval, and retries. Search millions of calls by customer, prompt version, or cost.

LLM-as-judge evals

The same scorer that sits in CI runs on a sample of production. LLM judge, code check, or human. Prompt versions that drop faithfulness get held.

Spend and latency alerts

Error rate, cost, and latency trip Slack the minute they cross your line. Spend caps halt LLM traffic at the limit you set.

AI features

Four jobs on model traffic

Each card names the AI job, who it is for, and the route. Glimrel is infrastructure for LLM and agent traffic. It does not train foundation models, and it does not act as an agent.

AI feature · Multi-model routing

One host across the families you already call

For platform teams wiring OpenAI, Anthropic, Gemini, Llama, Nova, or Groq

Keep the OpenAI-shaped client. Change the base URL and the key. Fallbacks, caches, and spend caps attach to that key. Planned access: Bedrock, Anthropic, OpenAI direct, Groq, Hugging Face.

SDK → Glimrel gateway → provider fallback → POST /v1/chat/completions

AI feature · Nested LLM traces

Tools, retrieval, and retries as spans

For engineers debugging agent and RAG traffic

Every completion, tool call, and retry is a span you can search. Customer, prompt version, and cost sit on the tree. This is observability for LLM traffic — not a chatbot.

Request → nested spans → GET /v1/traces/{id}

AI feature · LLM-as-judge evals

The CI grader scores a live sample

For quality leads who refuse silent drift

LLM judge, code check, or human. The same scorer that already sits in CI runs on a sample of production spans. Prompt versions that drop faithfulness get held.

Sample prod → same grader as CI → POST /v1/evals/runs

AI feature · Agent traffic plane

LangChain and Assistants sit on the gateway

For teams that already ship agents

Glimrel is the plane under agent frameworks (LangChain, OpenAI Assistants, homebrew). It does not choose tools or act as an agent. Spend caps halt traffic at the limit you set.

Agent SDK → one Glimrel key → route, trace, grade

Teams

Who holds the knobs

Platform

Give every product team one host, one key shape, and one place to debug a failed tool call.

Finance

Budgets attach to customer, project, or model. Cache hits show up next to spend, so repeats stop billing twice.

Quality

The grader in CI scores a live sample. Prompt versions that drop faithfulness get blocked before they own the river.