LLM gateway · traces · evals

One endpoint for model traffic.

Glimrel is a core AI/ML control plane: route LLM calls, nest traces for tools and retries, and grade live spans with the same evals you already run in CI. Point the SDK you have at one host. Glimrel is not an agent product — it sits under the agents you already ship.

  • Gateway multi-model routing
  • Traces nested LLM spans
  • Evals LLM-as-judge in prod
Claude 3.7 SonnetGPT-4oGPT-4o miniGemini 2.5Llama 3 70BAmazon NovaGroq BedrockAnthropicOpenAI directHugging Face Claude 3.7 SonnetGPT-4oGPT-4o miniGemini 2.5Llama 3 70BAmazon NovaGroq BedrockAnthropicOpenAI directHugging Face

AI features

Where models sit in Glimrel

Four named jobs on LLM traffic. Each card is a feature: what it does, who it is for, and the route. Glimrel routes, traces, and grades model calls. It does not train those models, and it does not ship agents.

AI feature · Multi-model routing

One host across the families you already call

For platform teams wiring OpenAI, Anthropic, Gemini, Llama, Nova, or Groq

Keep the OpenAI-shaped client. Change the base URL and the key. Fallbacks, caches, and spend caps attach to that key. Planned access paths: Bedrock, Anthropic, OpenAI direct, Groq, Hugging Face.

SDK → Glimrel gateway → provider fallback → POST /v1/chat/completions

AI feature · Nested LLM traces

Tools, retrieval, and retries as spans

For engineers debugging agent and RAG traffic

Every completion, tool call, and retry is a span you can search. Customer, prompt version, and cost sit on the tree. This is observability for LLM traffic — not a chatbot.

Request → nested spans → GET /v1/traces/{id}

AI feature · LLM-as-judge evals

The CI grader scores a live sample

For quality leads who refuse silent drift

LLM judge, code check, or human. The same scorer that already sits in CI runs on a sample of production spans. Prompt versions that drop faithfulness get held.

Sample prod → same grader as CI → POST /v1/evals/runs

AI feature · Agent traffic plane

LangChain and Assistants sit on the gateway

For teams that already ship agents

Glimrel is the plane under agent frameworks (LangChain, OpenAI Assistants, homebrew). It does not choose tools or act as an agent. Spend caps halt traffic at the limit you set.

Agent SDK → one Glimrel key → route, trace, grade

Model families on the route sheet: Claude 3.7 Sonnet, GPT-4o / GPT-4o mini, Gemini 2.5, Llama 3 70B, Amazon Nova, Groq. Glimrel does not train those models. MVP stage — no invented customers or uptime claims on this page.

Method

How Glimrel ships

Swap the host. Watch the current. Sample production with the grader you already trust.

01

Swap the base.

Keep your OpenAI client. Change the host and the key. Routing, fallbacks, and caches attach to that key.

02

Watch the current.

Requests, tokens, cache hits, and spend draw on one dashboard. Search nested LLM spans across millions of calls.

03

Grade the live set.

Sample production, score it with the LLM judge or code check you already run in CI, and hold the prompt version that drops below your line.

Stack

What you actually buy

An LLM control plane. Gateway, traces, and evals — not a chat widget and not a trained foundation model.

Gateway

One base URL, fallbacks, caches, and budgets per key, customer, or org.

Traces

Nested spans for tools, retrieval, and retries. Search millions of LLM calls.

Evals

LLM judge, code check, or human. Same scorer on a fixture and on 5% of prod.

Alerts

Error rate, cost, and latency trip Slack the minute they cross your line.

Start

A gateway key in one afternoon.

Write team@glimrel.store. You get a test token, a sample gateway project, and a 30-minute walkthrough on the LLM traffic you want to watch.

Request a key