Apply policies across services
Bring authentication, rate limits, and logging together for REST and gRPC APIs, LLM requests, and agent calls.
Explore the API gateway →One gateway for your APIs, LLMs and AI agents, with one set of identity, rate-limit, budget and guardrail policies. Built to run in your VPC or on-premises.
Tygress is pre-release: built and in active pre-1.0 development, not yet publicly available. Join for beta access and build updates. No payment required.
Built for platform teams
Tygress is for platform teams bringing APIs, models, and agents under shared security and cost policies, with a gateway they run in their own infrastructure.
Bring authentication, rate limits, and logging together for REST and gRPC APIs, LLM requests, and agent calls.
Explore the API gateway →Route LLM requests across providers with virtual keys, USD budgets, and failover policies in one place.
Explore the AI gateway →Built to run as one binary, or as separate control-plane and data-plane roles, from a container image or the Docker Compose reference stack, in your VPC or on-premises.
Explore deployment options →What's new
Tygress is pre-release, and the gateway is already built: 40 dated updates since June 2026, the latest on .
Latest: Failover errors that explain themselves ·
Multi-turn function calling over /v1/responses works on every provider whose model supports tools, buffered or streaming; providers without a native Responses API are bridged.
Tygress now runs on its own async transport and request engine, so guardrails can see every streamed chunk and each retry is translated, and SigV4-signed for Bedrock, over its own bytes.
Sandboxed proxy-wasm plugins get one guest per request and schema-checked config; HTTP callouts, shared data, queues and timers are not implemented yet.
Clients that speak Anthropic Messages reach every provider: native on Anthropic, bridged elsewhere, with DLP and the prompt guard on bridged traffic.
An internal two-round review of the proxy, auth, admin API, AI and MCP layers fixed 27 verified vulnerabilities before release.
Teams share a USD budget alongside each member's own, and one admin call mints a key with its budget, limits, allowed models and expiry.
Pre-1.0: the configuration schema and admin API may still change between releases.
API gateway
Routing, identity and traffic control for REST, gRPC and WebSocket services, plus TCP/UDP and TLS passthrough at L4, written in Rust on Tygress's own async transport. No enterprise-only features: planned pricing puts every gateway capability in the $0 tier.
Read about the self-hosted enterprise API gateway →HTTP/1.1 and HTTP/2 with ALPN and h2c, gRPC, gRPC-Web, and gRPC-JSON transcoding from descriptors loaded at runtime, with no codegen. Plus WebSocket, TCP/UDP/TLS passthrough and HTTP CONNECT forward proxying.
Trie-based routing on host, SNI, path, method, headers, query parameters, cookies and gRPC service or method. Match exactly, by regex, by prefix or by presence, and negate any condition. Split traffic by weight for canaries, or mirror a sample to a shadow upstream.
API keys, basic and HMAC auth, JWT with remote JWKS, JWE, OIDC, OAuth2 introspection, LDAP and Active Directory, mTLS and forward-auth, with ACL and OPA for authorization. They are built-in plugins, not enterprise add-ons.
Round-robin, weighted, least-request and consistent-hash load balancing with sticky sessions. Active HTTP, TCP and gRPC health checks, passive checks and outlier ejection. Retries with a retry budget, circuit breaking and concurrency limits.
Request limits with token bucket or fixed window, shared across replicas with Redis. Key a limit on the client IP, consumer, route, a header or a JWT claim, and cap requests in flight so overload is shed as 503s.
From bot detection, CSRF protection and IP restriction to response caching with HTTP PURGE and gzip, brotli or zstd compression. Extend with sandboxed proxy-wasm or inline Rhai/JavaScript, without rebuilding the gateway.
Declarative YAML or JSON config. tygress validate checks it offline, and the admin API's dry-run puts a change through the same checks as a real apply without committing it. Route, upstream, consumer and AI changes made through the admin API swap in atomically; in-flight requests finish on the config they started with.
Prometheus metrics, OpenTelemetry tracing and structured JSON logs, with log sinks for HTTP, TCP, UDP, syslog, files, Splunk HEC, Loki and Elasticsearch. tygress explain-chain prints the plugins each route runs, in order, and which scope's config wins.
No garbage-collection pauses, and the transport crate forbids unsafe code. proxy-wasm plugins run in a WebAssembly sandbox with memory, instruction and time limits.
AI gateway
Clients keep their OpenAI or Anthropic SDK and reach nine provider types through Chat Completions, Responses or Messages. Keys, budgets, guardrails and agent approvals run on the same gateway that already authenticates your API traffic.
Read about the self-hosted enterprise AI gateway →Teams share a USD budget alongside each member's own, with day, month or lifetime caps that PostgreSQL keeps across restarts. Windowed token and spend limits can also key on a route or an identity claim from your IdP. Spend is priced from your own rates or a bundled 80-model catalog, counting cached and reasoning tokens, and Slack-compatible alerts fire at 50, 80 and 95% by default.
Regex detectors with Luhn and entropy checks redact, block or audit PII and secrets in prompts, and can redact them in replies, streamed replies included. Redaction runs before provider translation and signing, so the provider receives the redacted text.
Heuristic rules block instruction overrides, role overrides and system-prompt probes, with look-alike Unicode letters folded before matching. An optional external classifier can fail open or closed. Replies, streamed ones included, are checked for system-prompt leaks.
Hold the MCP tool calls and A2A methods you mark as sensitive until an operator approves or denies them. With no decision, the request is denied after 45 seconds by default. Webhooks can notify Slack, Discord or Teams, and each decision is recorded in PostgreSQL when a database is configured.
Proxy MCP servers behind the same auth and rate limits as your APIs, with per-consumer tool allowlists on every tools/call. The gateway can also act as the MCP server itself, turning REST endpoints or an OpenAPI 3.x spec into MCP tools.
Route Agent2Agent calls by skill affinity or round-robin, with Agent Card discovery and one federated Agent Card for the agents behind the gateway. Automatic failover between agents is not wired yet.
A similar prompt reuses a stored answer when its embedding clears a cosine-similarity threshold you set. Answers live in memory, Redis or pgvector, and each lookup makes one embedding call, to an OpenAI-compatible endpoint or an opt-in local ONNX model. Streaming and tool-calling requests bypass the cache.
An opt-in control loop tunes weights across the providers you list, toward lower cost, lower latency or more resilience, within limits you set, with automatic rollback or human approval.
AI-specific Prometheus metrics on top of standard HTTP telemetry: tokens, cost, time to first token and cache lookups, plus GenAI OpenTelemetry spans and Langfuse export. Usage and spend roll up per consumer and per team through the admin API.
Providers
Eight vendor adapters, plus one for any endpoint that speaks the OpenAI API. Clients call any of them through any of three APIs: natively where the provider offers that API, bridged through chat completions where it doesn't.
POST /v1/chat/completions
All nine provider types
Passed through to each provider's OpenAI-compatible endpoint, including the compatibility layers Anthropic, Gemini and Vertex AI offer. Tygress translates only for Bedrock models on the Converse API. Vendor features a compatibility layer doesn't carry need the vendor's native API.
POST /v1/responses
Native on OpenAI and Azure OpenAI
Bridged through chat completions on the other seven provider types, with function calling, buffered or streaming.
POST /v1/messages
Native on Anthropic
Bridged through chat completions on the other eight provider types, with tool calling, buffered or streaming.
See the LLM API compatibility table by provider type →
Consumers get gateway keys bound to provider keys the operator holds, so raw provider keys stay in the gateway. One admin call issues a key with a budget, request and token limits, allowed models and an expiry.
Pool several API keys per provider; a key that gets a 429 rests while the others serve. Rate-limit headers from OpenAI, Azure OpenAI, Anthropic and Groq drive quota-aware rotation. Fail over across providers, with each target's failure reason listed in the 502.
Route by model, or by fallback, round-robin, weighted or semantic selection. Semantic routing matches each prompt to a complexity tier so simpler prompts can go to cheaper models, and runs in shadow mode until you enforce it.
Engine
Since September 2026, Tygress runs on its own async transport and request engine, written in Rust. Owning the transport changed what the gateway can do with each request.
Provider translation and request signing run inside each retry and failover attempt, after plugins edit the request. A signed request, such as SigV4 to Bedrock, covers the exact bytes that attempt sends, DLP redactions included.
With reply scanning on, DLP and the prompt guard's leak check inspect each chunk of a streamed SSE reply as it passes, not one buffered body. DLP matches in replies are redacted before they reach the client.
tygress --upgrade hands the listening sockets to the new binary over a Unix socket, on Linux and macOS, while the old process drains in-flight work. On SIGTERM, /healthz returns 503 so load balancers deregister the node before it stops accepting.
HTTP/1.1, HTTP/2, gRPC-Web, WebSocket, HTTP CONNECT forward proxying and L4 TCP/UDP/TLS passthrough run on the same transport, so one drain and upgrade path covers all of them.
5,800+
automated tests
Test functions across unit, integration, and 29 end-to-end suites, run in CI.
Daily
dependency-advisory audit
cargo-deny checks dependencies against RustSec advisories every day; advisory fixes have shipped within days.
Rust 1.94
minimum Rust version, checked in CI
CI also runs lint, tests, a release build, WebAssembly guest builds and Docker-backed suites.
27
verified vulnerabilities fixed
An internal two-round security review in July 2026, followed by a September hardening pass. Security review in the changelog →
Benchmark figures shown here earlier measured the previous engine and have been withdrawn. New results will be published with their methodology.
Read the changelog entry on the new transport →Enterprise
A gateway built for your platform team to self-host, with the identity, audit and availability controls security teams ask for. None of them is held back for a paid edition: the planned pricing puts every gateway capability in the $0 tier.
A Dockerfile builds a distroless, non-root container image, and a Docker Compose reference stack runs a control plane, two data planes and NATS. No image is published yet. An opt-in build discovers upstream pods through Kubernetes EndpointSlices. Requests to cloud model providers still leave your network, so keeping prompts local needs local model endpoints.
OIDC, JWT with remote JWKS, LDAP/Active Directory, mutual TLS, OAuth2 introspection with required scopes, JWE and HMAC, combined per route and mapped to a consumer. ACL and OPA handle authorization.
The control plane writes each admin change as a versioned snapshot to NATS JetStream. Data planes follow it, rebuild atomically and keep a last-known-good copy on disk. Redis shares rate-limit counters and approval queues across replicas.
Human approvals hold MCP tool calls and A2A methods at the gateway; unanswered requests are denied after a timeout (45 seconds by default). With PostgreSQL configured, approval decisions, autopilot changes and control-plane config revisions are kept as append-only records.
Attach rate limits, model allowlists and MCP tool allowlists to a consumer or group, key per-user limits on IdP claims, and give each team a shared USD budget alongside per-member budgets. Usage APIs report per consumer and per team.
Admin API tokens and the operator console (an opt-in build) use admin or read-only roles, and the console supports OIDC single sign-on. The incident-response plugin can ban flooding or abusive client IPs and open circuits on upstream outages, lifting each remediation when it expires.
Post-quantum TLS
Adversaries can record encrypted traffic today and decrypt it once quantum computers mature. To blunt these “harvest-now, decrypt-later” attacks, Tygress offers the hybrid X25519MLKEM768 group first on its TLS listeners and upstream connections: ML-KEM-768 (NIST FIPS 203) combined with X25519. Peers that support it negotiate it with no configuration; classical-only peers fall back to classical key exchange.
Key exchange only: certificates and signatures remain classical. Post-quantum TLS in the changelog →
Compare
Many enterprises run an API gateway, a separate LLM proxy, and custom agent middleware, then synchronize policy across all three. Tygress brings those jobs into one self-hosted gateway, with one set of identity, rate-limit, budget and guardrail policies.
| Capability | Tygress | Kong | Apigee | LiteLLM | Portkey (Palo Alto Networks) |
|---|---|---|---|---|---|
| Availability | Pre-release | Generally available | Generally available | Generally available | Generally available |
| API gateway | Pass-through only | ||||
| AI / LLM gateway | Built in | Separate AI Gateway 2.x | Managed | Native | Native |
| Gateway-held approvals | MCP + A2A, default deny | Not documented | Not documented | Client-side | Not documented |
| Enterprise identity (OIDC/LDAP/mTLS) | OIDC/mTLS: Enterprise | OIDC/mTLS (no LDAP) | Partial | Partial | |
| MCP gateway | Built in | AI license | Managed | Native | |
| A2A agent gateway | Built in | AI license | Not documented | Agent Gateway | |
| Self-hosted / air-gapped | Yes (AI 2.x needs Konnect) | Hybrid; Private Cloud air-gap | Yes (air-gap: Enterprise) | OSS + hybrid; no air-gap | |
| AI governance & DLP | Built in | AI plugins | Managed | Built in + integrations | Partial |
| Multi-provider failover | AI license | Managed | |||
| Semantic caching | AI license | Managed | Via cache backend | Paid tiers | |
| Post-quantum TLS | Hybrid key exchange by default; per-route enforcement; no PQ certificates | Hybrid key exchange + ML-DSA certificates (non-FIPS) | Client leg via Cloud LB | Not documented | Not documented |
| Data-plane runtime | Rust (own transport) | NGINX / Lua | Managed / hybrid | Python + Rust (beta) | Node/TS |
| Open source | Core planned at launch | Limited | Yes | Partial |
Tygress column reflects the current pre-release build. Competitor columns were checked against public documentation on 6 October 2026. “AI license” marks Kong features its documentation labels “AI License Required”. Portkey is now part of Palo Alto Networks as Prisma AIRS AI Gateway. Product editions and competitor features change; verify requirements during procurement.
Pricing
Planned pricing. Paid plans add managed hosting, support, SLAs and services, not gateway features.
Self-hosted
$0
The whole gateway on your own infrastructure: every capability on this page.
Join waitlistManaged by Tygress
Tygress runs and upgrades the gateway for you.
Your infrastructure
Data planes in your cloud account; the control plane managed by Tygress.
Services and SLAs
For teams running Tygress on-premises or air-gapped who need a support contract.
All prices are preliminary and may change before launch. An open-source core is planned at launch. Release and license details will be published with the release.
FAQ
Availability, current capabilities, and how Tygress fits your API and AI workloads.
Tygress is a self-hosted API gateway and AI gateway written in Rust on its own purpose-built async transport. It routes REST, gRPC and WebSocket APIs, LLM requests, MCP tools and A2A agents through one data plane with shared identity, rate limiting, cost and guardrail policies, and also proxies TCP/UDP traffic at L4.
Tygress is pre-release. The gateway is built and in active pre-1.0 development, but it is not publicly available yet. Join the early-access waitlist for beta availability and product updates.
Yes, if it understands both. An API gateway manages identity and traffic for REST and gRPC services; an AI gateway adds model routing, provider credentials, token budgets, and AI guardrails. Tygress runs both in one self-hosted data plane, so one identity, rate-limit, and logging model covers APIs, models, and agents. Whether you should consolidate depends on your workloads and existing infrastructure.
No. Earlier builds used a vendored fork of Cloudflare's Pingora. Since September 2026, Tygress runs on tygress-proxy, its own async transport, and its own request engine, written in Rust. Owning the transport lets DLP and prompt guards inspect every chunk of a streamed reply, and runs provider translation and request signing inside each retry and failover attempt, after every plugin edit.
Tygress runs as one binary, or as separate control-plane and data-plane roles that sync configuration over NATS. The source includes a Dockerfile for a distroless, non-root container image and a Docker Compose reference stack; no image is published yet. An opt-in Kubernetes build watches EndpointSlices to track upstream pods. It is designed to run in a private VPC or on-premises; fully air-gapped operation also requires local model endpoints, identity services, and other dependencies.
Self-hosting the gateway gives you control of routing and policy. Requests sent to external model APIs still leave your network; keeping inference local requires local model endpoints and control of logs, caches, and telemetry.
Tygress connects to OpenAI, Anthropic, Google Gemini, Google Vertex AI (Gemini models), Azure OpenAI, AWS Bedrock, Mistral, Groq, and any OpenAI-compatible endpoint. Clients can call any of them through the OpenAI Chat Completions, OpenAI Responses, or Anthropic Messages API: natively where the provider supports that API, and bridged through chat completions otherwise. Chat Completions reaches Anthropic, Gemini, and Vertex through each vendor's own OpenAI-compatible endpoint. Some provider-specific features, such as stored conversations, hosted tools, prompt caching, or image input, need a native endpoint.
Tygress reads token usage from provider replies, including cached and reasoning tokens where the provider reports them, and prices it from a bundled 80-model catalog or your own pricing. Token limits and USD budgets run over fixed, sliding, rolling, or token-bucket windows keyed per consumer, route, or any request variable. Durable daily, monthly, or lifetime caps apply per consumer and per team, with webhook alerts at thresholds you set. Semantic caching can answer a similar non-streaming prompt without calling the chat model, and semantic routing, once switched from shadow mode to enforce, can send simpler prompts to cheaper models. Actual savings depend on your models, traffic, and cache settings.
Authentication plugins cover OIDC, JWT with remote JWKS, LDAP and Active Directory, mTLS, OAuth2 token introspection, JWE, HMAC, basic auth, and API keys, with OPA and ACLs for authorization. AI controls include PII and secret redaction, a prompt-injection guard, model and MCP tool allowlists, and human approvals for MCP tool calls and A2A methods, with each decision recorded in an append-only PostgreSQL table when a database is configured. These controls support your security policies; deployment choices and upstream services also affect how data is handled.
Tygress is pre-1.0, and some things are not done. There is no published container image, Helm chart, or Kubernetes manifest, and the operator console, Kubernetes discovery, Kafka logging, and the local embedder for the semantic cache are opt-in builds. HTTP/3, WebSocket over HTTP/2, SAML, post-quantum certificates and signatures, and Windows are not supported. The OpenAI Realtime API returns 501, failover between A2A agents is not wired, and some proxy-wasm hostcalls (HTTP callouts, shared data and queues, timers) are not implemented. Benchmarks will be published with their methodology. The configuration schema and admin API may change before 1.0.
An open-source core is planned at launch. Release and license details will be published with the release.
Current capabilities include: API gateway for HTTP/1.1, HTTP/2, gRPC (including gRPC-Web and JSON transcoding), WebSocket, and TCP/UDP/TLS passthrough, with load balancing, active and passive health checks, retries, and circuit breaking; 60+ built-in plugins for authentication (OIDC, JWT, LDAP, mTLS, OAuth2 introspection), authorization (ACL, OPA), traffic control, and logging, plus sandboxed proxy-wasm plugins and inline Rhai or JavaScript; LLM routing across nine provider types with failover, virtual keys, token limits, USD budgets for keys and teams, and a bundled model-pricing catalog; AI guardrails: prompt-injection detection, DLP that redacts, blocks, or audits PII and secrets in prompts and can redact them in buffered and streamed replies, model allowlists, and semantic caching; MCP tool gating and human approvals, REST-to-MCP toolsets, and A2A agent routing; Prometheus metrics, OpenTelemetry tracing with GenAI semantic conventions, Langfuse export, and append-only PostgreSQL records of config revisions, approval decisions, and autopilot changes; Hybrid post-quantum TLS key exchange (X25519MLKEM768), offered by default and enforceable per listener, upstream, or route; ACME certificates; and zero-downtime binary upgrades.
Related guide: AI gateway vs API gateway
Evaluating a gateway for your APIs or AI workloads? Tygress is pre-release. The gateway is built and in active pre-1.0 development, but it is not publicly available yet. Join the early-access waitlist for beta availability and product updates.