Requesty

The AI Gateway for Production

Access 600+ models through one API with intelligent routing, real-time analytics and centralized governance.

Or email us at sales@requesty.ai

app.requesty.ai/analytics
R
Requesty
Balance
+ Top Up
$4,182.14
OBSERVABILITY
Analytics
Leaderboard
Logs
Audit Logs
AI GATEWAY
Model Library
Model Discovery
MCP Gateway
CONFIGURATIONS
Prompt Library
API Keys
Routing Policies
Bring Your Own Keys
My Groups
ADMIN
Admin Panel
Settings
T
Thibault Jaigu

Analytics

7 Days30 DaysThis Month
General
Advanced
Security
Cost Overview$445.95
$0$36$73Feb 24Feb 25Feb 26Feb 27Feb 28Mar 1Mar 2Mar 3
Request Volume123,200
09.8K19.7KFeb 24Feb 25Feb 26Feb 27Feb 28Mar 1Mar 2Mar 3
Token Usage30.4M
02530.0K5060.0KFeb 24Feb 25Feb 26Feb 27Feb 28Mar 1Mar 2Mar 3
Total Request Latency407ms
340ms415ms490msFeb 24Feb 25Feb 26Feb 27Feb 28Mar 1Mar 2Mar 3
Model Usage5 models
ModelRequestsTokensCostAvg Latency
opus-542,180(34.2%)18.4M(38.0%)$163.15892ms
gpt-5.6-sol31,420(25.5%)12.8M(26.4%)$114.70445ms
gemini-3.7-flash24,100(19.6%)9.6M(19.8%)$81.20512ms
deepseek-v4-flash15,600(12.7%)5.2M(10.7%)$52.90380ms
glm-5.39,900(8.0%)2.4M(5.0%)$34.00310ms
Trusted by teams at
Shopify
Amadeus
Chargebee
Contentful
Demandbase
Pfizer
PWC
Capgemini
Sage
Siemens
Relevance AI
Appnovation
Shopify
Amadeus
Chargebee
Contentful
Demandbase
Pfizer
PWC
Capgemini
Sage
Siemens
Relevance AI
Appnovation
600+
Models
99.99%
Uptime
<14ms
Failover
225B+
Tokens/day

Your AI ops, at a glance

Heatmaps, cost breakdowns, cache gauges across all your AI providers.

Total Requests
143.2K
+18.7%vs last month

Across all providers and models. Peak: 8.2K/hr at 14:00 UTC.

Total Cost
$1,247
-12.3%vs last month

Down from $1,422 last month thanks to semantic caching and smart routing.

Cache Hit Rate
78.6%
+4.2%vs last month

111.8K cache hits saved $978 this month. Semantic matching enabled.

Uptime
99.97%
30 days

Auto-failover triggered 3 times. Zero downtime for your users.

Latency Distribution
low
med
high
<50ms
<200ms
<500ms
<1s
<2s
>2s
0003060912151821
Cache Performance
78.6%
Hit Rate
111.8K
Hits
30.4K
Misses
$978
Saved
Daily Cost by Model
7 days30 days90 days
MonTueWedThuFriSatSun
opus-5
gpt-5.6-sol
gemini-3.7-flash
deepseek-v4-flash
glm-5.3
By Model
opus-5
$42634.2%
gpt-5.6-sol
$31825.5%
gemini-3.7-flash
$24419.6%
deepseek-v4-flash
$15912.7%
glm-5.3
$1008%
Total 143.2K requests$1,247
from openai import OpenAI
client = OpenAI(
base_url="https://router.requesty.ai/v1",
api_key="your-requesty-key"
)
response = client.chat.completions.create(
model="anthropic/claude-opus-5",
messages=[{"role": "user", "content": "Hello!"}]
)
print(response.choices[0].message.content)
Plug & Play

Integrate in
a minute

Integrate Requesty in just 3 lines of code. No changes to your existing stack. Use the OpenAI SDK you already know.

OpenAI-compatible API, works with any SDK
No vendor lock-in. Switch models with one line
Automatic failover & load balancing included
Infrastructure

Route every request

Pin traffic to a region, cap what each team can spend, and fail over to a healthy provider without changing a line of code.

Geo-Based Routing

Route requests to the nearest region automatically. EU data stays in Frankfurt, US in Virginia, APAC in Singapore. Full data residency compliance.

Policy-Based Controls

Set spending limits, model restrictions, and rate limits per user, team, or API key. Policies cascade from organization to individual level.

Automatic Failover

When a provider goes down, traffic switches to the next best option in under 14ms. Zero downtime, zero manual intervention.

Agent Routing Policies

Define routing strategies per agent. Assign preferred models, fallback chains, and cost caps so each agent gets the right model for the job.

Request Routing
Live
eu-central-1
Frankfurt
Primary
Latency: 12ms
Requests: 14.2K/min
us-east-1
Virginia
Active
Latency: 8ms
Requests: 22.1K/min
ap-southeast-1
Singapore
Active
Latency: 18ms
Requests: 6.8K/min
Data Residency
EU requests → eu-central-1 only
Observability

See where the money goes

Track spend by model, team and agent, catch a latency regression the hour it starts, and find the call that caused it.

Cost Analytics

Track spending by model, user, and team in real-time. See exactly where every dollar goes across all your AI providers.

Performance Monitoring

Monitor latency, success rates, and token usage across every provider. Get alerted before issues impact your users.

Usage Insights

Understand which models, teams, and users drive consumption. Make data-driven decisions about your AI infrastructure.

Agent Analytics

Track latency, cost, and success rates per agent. See which agents perform best and where bottlenecks hide.

Cost by Model
7 days30 days
Cost
$1,247
Requests
143.2K
Tokens
48.2M
Avg Latency
412ms
M
T
W
T
F
S
S
opus-5
gpt-5.6-sol
gemini-3.7-flash
deepseek-v4-flash
glm-5.3
Governance

Define and enforce the rules

Strip PII before it reaches a provider, block what your policy forbids, decide what each team can use, and hand auditors a log of who changed what.

PII Detection & Scrubbing

Automatically detect and redact personal data before it reaches the model. Emails, phone numbers, SSNs, credit cards, all scrubbed in real-time.

Content Guardrails

Enforce content policies, block prompt injections, and filter harmful outputs. Protect your users and your brand automatically.

Team Management

Role-based access with Owner, Admin, Developer, and Viewer roles. Set per-team budgets, model allowlists, and usage quotas.

Audit Logs

Complete audit trail of every action. Track who did what, when, and from where. Export logs for compliance and forensics.

PII Scanner
Active
Incoming Request
"Please help john.doe@acme.com with account #4521-8834-1290"
Detected
EMAIL
ACCOUNT_ID
Scrubbed Output
"Please help [EMAIL] with account [ACCOUNT_ID]"
2 entities detected • Scrubbed in 3ms
Review and residency

Built to survive procurement

The questions security and legal ask, answered before they ask them.

In progress

SOC 2 Type II

Observation period under way with an independent auditor. Controls in place and documented; the target date is on the trust page.

Contractual

GDPR, DPA on request

An Article 28 DPA, signed on request at any spend, with sub-processors listed.

Regional

EU residency

Frankfurt, on AWS eu-central-1, via router.eu.requesty.ai.

Retention

Zero data retention, pinnable

Pin your traffic to endpoints that retain nothing. 131 of the EU model endpoints qualify.

Pricing

One number, and it sits on top of the model cost

No per-seat pricing, no minimum spend. You pay for what your application spends, plus 5%.

Pay as you go
5%markup on model cost

Every feature included: routing, caching, failover, guardrails, analytics and audit logs. No seat fees, no feature tiers.

See pricing
0%
Bring your own keys
on your own provider contracts

Keep the pricing you have negotiated with each provider and still get routing, limits and observability across all of them.

How BYOK works
Volume
Enterprise
discounts and committed spend

SSO, SCIM, private regions, a signed DPA and a named contact. Procurement paperwork handled by people, not a form.

Speak to a founder

Questions & Answers

An AI gateway between your app and 600+ LLM providers. Change your base URL to router.requesty.ai and instantly get intelligent routing, fallbacks, cost optimization, caching, governance, and observability.

One line of code: client = OpenAI(base_url='https://router.requesty.ai/v1', api_key='your-key'). Works with all major SDKs.

Smart routing to cheaper equivalent models, caching, automatic fallback from expensive providers, per-user spending limits, and real-time cost analytics.

Yes. Native OAuth integrations. Any model, unlimited requests, no rate limits.

5% markup on model costs. All features included. Enterprise plans available with volume discounts.

Yes. Bring your own keys for any provider while getting Requesty's routing and observability. Or use our unified key.

Start building with Requesty

One line of code. 600+ models. Full control.

Speak to founders