Pre-release · 40 updates since June

Self-hosted API gateway & AI gateway, in Rust

One gateway for your APIs, LLMs and AI agents, with one set of identity, rate-limit, budget and guardrail policies. Built to run in your VPC or on-premises.

Tygress is pre-release: built and in active pre-1.0 development, not yet publicly available. Join for beta access and build updates. No payment required.

One binaryall upstreams healthy
Ingress
GET /api/v1/users
POST /v1/messages
POST /v1/responses
MCP tools/call
tygress
one policy chain
auth · limits · budgets · guardrails
Upstreams
REST · gRPC
OpenAI
Anthropic
Bedrock
Vertex AI

Built for platform teams

One policy layer for APIs and AI

Tygress is for platform teams bringing APIs, models, and agents under shared security and cost policies, with a gateway they run in their own infrastructure.

At a glance

Status
Pre-release: pre-1.0, not yet publicly available
Open source
Core planned at launch
Runs on
Linux and macOS
Deploys as
One binary, or control-plane and data-plane roles that sync over NATS
Optional dependencies
PostgreSQL (durable budgets, usage, config and approval records) and Redis (state shared across replicas)

Apply policies across services

Bring authentication, rate limits, and logging together for REST and gRPC APIs, LLM requests, and agent calls.

Explore the API gateway →

Control model access and spend

Route LLM requests across providers with virtual keys, USD budgets, and failover policies in one place.

Explore the AI gateway →

Self-hosted by design

Built to run as one binary, or as separate control-plane and data-plane roles, from a container image or the Docker Compose reference stack, in your VPC or on-premises.

Explore deployment options →

What's new

Built, not just planned

Tygress is pre-release, and the gateway is already built: 40 dated updates since June 2026, the latest on .

built-in plugins
60+
LLM provider types
9
LLM APIs on every provider1
3
automated tests2
5,800+
  1. Chat Completions, Responses, and Messages: native where the provider offers the API, bridged through chat completions otherwise. Bridges refuse features they can't carry, such as stored conversations, hosted tools, or image input, instead of dropping them.
  2. Test functions across unit, integration, and 29 end-to-end suites, run in CI.

Latest: Failover errors that explain themselves ·

AI gatewayAgents & MCP

Function calling on /v1/responses, on every provider

Multi-turn function calling over /v1/responses works on every provider whose model supports tools, buffered or streaming; providers without a native Responses API are bridged.

Engine

Tygress now runs on its own transport and request engine

Tygress now runs on its own async transport and request engine, so guardrails can see every streamed chunk and each retry is translated, and SigV4-signed for Bedrock, over its own bytes.

API gatewayEngine

proxy-wasm plugins that keep per-request state

Sandboxed proxy-wasm plugins get one guest per request and schema-checked config; HTTP callouts, shared data, queues and timers are not implemented yet.

Read the full changelog →

Pre-1.0: the configuration schema and admin API may still change between releases.

API gateway

The API gateway underneath: protocols, identity, traffic

Routing, identity and traffic control for REST, gRPC and WebSocket services, plus TCP/UDP and TLS passthrough at L4, written in Rust on Tygress's own async transport. No enterprise-only features: planned pricing puts every gateway capability in the $0 tier.

Read about the self-hosted enterprise API gateway →

HTTP, gRPC, WebSocket and L4

HTTP/1.1 and HTTP/2 with ALPN and h2c, gRPC, gRPC-Web, and gRPC-JSON transcoding from descriptors loaded at runtime, with no codegen. Plus WebSocket, TCP/UDP/TLS passthrough and HTTP CONNECT forward proxying.

Routing on host, path, headers and more

Trie-based routing on host, SNI, path, method, headers, query parameters, cookies and gRPC service or method. Match exactly, by regex, by prefix or by presence, and negate any condition. Split traffic by weight for canaries, or mirror a sample to a shadow upstream.

Identity built in

API keys, basic and HMAC auth, JWT with remote JWKS, JWE, OIDC, OAuth2 introspection, LDAP and Active Directory, mTLS and forward-auth, with ACL and OPA for authorization. They are built-in plugins, not enterprise add-ons.

Resilience

Round-robin, weighted, least-request and consistent-hash load balancing with sticky sessions. Active HTTP, TCP and gRPC health checks, passive checks and outlier ejection. Retries with a retry budget, circuit breaking and concurrency limits.

Rate limiting

Request limits with token bucket or fixed window, shared across replicas with Redis. Key a limit on the client IP, consumer, route, a header or a JWT claim, and cap requests in flight so overload is shed as 503s.

60+ built-in plugins

From bot detection, CSRF protection and IP restriction to response caching with HTTP PURGE and gzip, brotli or zstd compression. Extend with sandboxed proxy-wasm or inline Rhai/JavaScript, without rebuilding the gateway.

Validate before you apply

Declarative YAML or JSON config. tygress validate checks it offline, and the admin API's dry-run puts a change through the same checks as a real apply without committing it. Route, upstream, consumer and AI changes made through the admin API swap in atomically; in-flight requests finish on the config they started with.

Observability

Prometheus metrics, OpenTelemetry tracing and structured JSON logs, with log sinks for HTTP, TCP, UDP, syslog, files, Splunk HEC, Loki and Elasticsearch. tygress explain-chain prints the plugins each route runs, in order, and which scope's config wins.

Written in Rust

No garbage-collection pauses, and the transport crate forbids unsafe code. proxy-wasm plugins run in a WebAssembly sandbox with memory, instruction and time limits.

AI gateway

LLMs, MCP tools and agents under one set of policies

Clients keep their OpenAI or Anthropic SDK and reach nine provider types through Chat Completions, Responses or Messages. Keys, budgets, guardrails and agent approvals run on the same gateway that already authenticates your API traffic.

Read about the self-hosted enterprise AI gateway →

Budgets that follow your org chart

Teams share a USD budget alongside each member's own, with day, month or lifetime caps that PostgreSQL keeps across restarts. Windowed token and spend limits can also key on a route or an identity claim from your IdP. Spend is priced from your own rates or a bundled 80-model catalog, counting cached and reasoning tokens, and Slack-compatible alerts fire at 50, 80 and 95% by default.

AI DLP and PII redaction

Regex detectors with Luhn and entropy checks redact, block or audit PII and secrets in prompts, and can redact them in replies, streamed replies included. Redaction runs before provider translation and signing, so the provider receives the redacted text.

Prompt injection guard

Heuristic rules block instruction overrides, role overrides and system-prompt probes, with look-alike Unicode letters folded before matching. An optional external classifier can fail open or closed. Replies, streamed ones included, are checked for system-prompt leaks.

Human approvals for agent actions

Hold the MCP tool calls and A2A methods you mark as sensitive until an operator approves or denies them. With no decision, the request is denied after 45 seconds by default. Webhooks can notify Slack, Discord or Teams, and each decision is recorded in PostgreSQL when a database is configured.

MCP gateway

Proxy MCP servers behind the same auth and rate limits as your APIs, with per-consumer tool allowlists on every tools/call. The gateway can also act as the MCP server itself, turning REST endpoints or an OpenAPI 3.x spec into MCP tools.

A2A agent routing

Route Agent2Agent calls by skill affinity or round-robin, with Agent Card discovery and one federated Agent Card for the agents behind the gateway. Automatic failover between agents is not wired yet.

Semantic caching

A similar prompt reuses a stored answer when its embedding clears a cosine-similarity threshold you set. Answers live in memory, Redis or pgvector, and each lookup makes one embedding call, to an OpenAI-compatible endpoint or an opt-in local ONNX model. Streaming and tool-calling requests bypass the cache.

Autopilot routing (opt-in)

An opt-in control loop tunes weights across the providers you list, toward lower cost, lower latency or more resilience, within limits you set, with automatic rollback or human approval.

AI observability

AI-specific Prometheus metrics on top of standard HTTP telemetry: tokens, cost, time to first token and cache lookups, plus GenAI OpenTelemetry spans and Langfuse export. Usage and spend roll up per consumer and per team through the admin API.

Providers

Nine provider types, three client APIs

Eight vendor adapters, plus one for any endpoint that speaks the OpenAI API. Clients call any of them through any of three APIs: natively where the provider offers that API, bridged through chat completions where it doesn't.

  • OpenAI
  • Anthropic
  • Google Gemini
  • Google Vertex AI · Gemini models
  • Azure OpenAI
  • AWS Bedrock
  • Mistral
  • Groq
  • + any OpenAI-compatible endpoint

POST /v1/chat/completions

OpenAI Chat Completions

All nine provider types

Passed through to each provider's OpenAI-compatible endpoint, including the compatibility layers Anthropic, Gemini and Vertex AI offer. Tygress translates only for Bedrock models on the Converse API. Vendor features a compatibility layer doesn't carry need the vendor's native API.

POST /v1/responses

OpenAI Responses

Native on OpenAI and Azure OpenAI

Bridged through chat completions on the other seven provider types, with function calling, buffered or streaming.

POST /v1/messages

Anthropic Messages

Native on Anthropic

Bridged through chat completions on the other eight provider types, with tool calling, buffered or streaming.

Bridges
Refuse what they can't carry instead of dropping it. Responses: stored conversations, hosted tools, reasoning items and multimodal input. Messages: images, documents, extended thinking and vendor tools.
Embeddings
OpenAI, Azure OpenAI, Gemini, Mistral and Bedrock, plus OpenAI-compatible endpoints where enabled. Not Anthropic, Groq or Vertex AI. Gemini embeddings and Cohere embeddings on Bedrock report no token counts, so they are not metered.
Other endpoints
Images, audio, moderations, files and batches work on the OpenAI provider type only, without cost accounting. The Realtime API returns 501.

See the LLM API compatibility table by provider type →

Virtual keys

Consumers get gateway keys bound to provider keys the operator holds, so raw provider keys stay in the gateway. One admin call issues a key with a budget, request and token limits, allowed models and an expiry.

Key pools and failover

Pool several API keys per provider; a key that gets a 429 rests while the others serve. Rate-limit headers from OpenAI, Azure OpenAI, Anthropic and Groq drive quota-aware rotation. Fail over across providers, with each target's failure reason listed in the 502.

Routing strategies

Route by model, or by fallback, round-robin, weighted or semantic selection. Semantic routing matches each prompt to a complexity tier so simpler prompts can go to cheaper models, and runs in shadow mode until you enforce it.

Engine

Why Tygress runs on its own transport

Since September 2026, Tygress runs on its own async transport and request engine, written in Rust. Owning the transport changed what the gateway can do with each request.

Transport
tygress-proxy, written for this gateway
HTTP framing
hyper and h2
TLS
rustls on aws-lc-rs
Unsafe code
Forbidden in the transport crate

Each attempt translated and signed

Provider translation and request signing run inside each retry and failover attempt, after plugins edit the request. A signed request, such as SigV4 to Bedrock, covers the exact bytes that attempt sends, DLP redactions included.

Guardrails on every streamed chunk

With reply scanning on, DLP and the prompt guard's leak check inspect each chunk of a streamed SSE reply as it passes, not one buffered body. DLP matches in replies are redacted before they reach the client.

Upgrades without closing listeners

tygress --upgrade hands the listening sockets to the new binary over a Unix socket, on Linux and macOS, while the old process drains in-flight work. On SIGTERM, /healthz returns 503 so load balancers deregister the node before it stops accepting.

One transport for HTTP, gRPC and L4

HTTP/1.1, HTTP/2, gRPC-Web, WebSocket, HTTP CONNECT forward proxying and L4 TCP/UDP/TLS passthrough run on the same transport, so one drain and upgrade path covers all of them.

How we check it

  • 5,800+

    automated tests

    Test functions across unit, integration, and 29 end-to-end suites, run in CI.

  • Daily

    dependency-advisory audit

    cargo-deny checks dependencies against RustSec advisories every day; advisory fixes have shipped within days.

  • Rust 1.94

    minimum Rust version, checked in CI

    CI also runs lint, tests, a release build, WebAssembly guest builds and Docker-backed suites.

  • 27

    verified vulnerabilities fixed

    An internal two-round security review in July 2026, followed by a September hardening pass. Security review in the changelog →

Benchmark figures shown here earlier measured the previous engine and have been withdrawn. New results will be published with their methodology.

Read the changelog entry on the new transport →

Enterprise

Identity, audit and uptime, on your infrastructure

A gateway built for your platform team to self-host, with the identity, audit and availability controls security teams ask for. None of them is held back for a paid edition: the planned pricing puts every gateway capability in the $0 tier.

Self-hosted in your VPC or on-premises

A Dockerfile builds a distroless, non-root container image, and a Docker Compose reference stack runs a control plane, two data planes and NATS. No image is published yet. An opt-in build discovers upstream pods through Kubernetes EndpointSlices. Requests to cloud model providers still leave your network, so keeping prompts local needs local model endpoints.

Enterprise identity, built in

OIDC, JWT with remote JWKS, LDAP/Active Directory, mutual TLS, OAuth2 introspection with required scopes, JWE and HMAC, combined per route and mapped to a consumer. ACL and OPA handle authorization.

Split control and data planes

The control plane writes each admin change as a versioned snapshot to NATS JetStream. Data planes follow it, rebuild atomically and keep a last-known-good copy on disk. Redis shares rate-limit counters and approval queues across replicas.

Approvals and audit records

Human approvals hold MCP tool calls and A2A methods at the gateway; unanswered requests are denied after a timeout (45 seconds by default). With PostgreSQL configured, approval decisions, autopilot changes and control-plane config revisions are kept as append-only records.

Consumers, groups and teams

Attach rate limits, model allowlists and MCP tool allowlists to a consumer or group, key per-user limits on IdP claims, and give each team a shared USD budget alongside per-member budgets. Usage APIs report per consumer and per team.

Admin roles and incident response

Admin API tokens and the operator console (an opt-in build) use admin or read-only roles, and the console supports OIDC single sign-on. The incident-response plugin can ban flooding or abusive client IPs and open circuits on upstream outages, lifting each remediation when it expires.

Post-quantum TLS

Post-quantum key exchange, offered by default and enforceable

Adversaries can record encrypted traffic today and decrypt it once quantum computers mature. To blunt these “harvest-now, decrypt-later” attacks, Tygress offers the hybrid X25519MLKEM768 group first on its TLS listeners and upstream connections: ML-KEM-768 (NIST FIPS 203) combined with X25519. Peers that support it negotiate it with no configuration; classical-only peers fall back to classical key exchange.

  • A policy per listener and per upstream — set pqc_policy to observe (the default), prefer or require. Require accepts only the hybrid group over TLS 1.3, so a classical-only peer fails the handshake instead of downgrading.
  • Enforcement per route or consumer — the pqc-guard plugin rejects requests whose handshake used classical key exchange on the routes or consumers you choose, or only warns or counts while you measure coverage.
  • Metrics on what clients negotiate — with a listener policy set, Prometheus counts inbound TLS requests by the key-exchange group each client negotiated.
  • A crypto-posture report — an admin endpoint classifies the certificates, key-exchange policies and token and auth algorithms the gateway is configured with, and maps them to published PQC-migration deadlines and PCI DSS 4.0.1 Req 12.3.3. Cipher suites are not part of the report.

Key exchange only: certificates and signatures remain classical. Post-quantum TLS in the changelog →

TLS 1.3 handshakehybrid PQ key exchange
ClientHello → key_share
X25519MLKEM768hybrid PQC
ServerHello ← ML-KEM ciphertext · ✓ established
classical-only peer → classical group, unless pqc_policy: require
FIPS 203
ML-KEM-768
HNDL
mitigated for PQ-capable peers
Posture
report

Compare

One gateway instead of a whole stack

Many enterprises run an API gateway, a separate LLM proxy, and custom agent middleware, then synchronize policy across all three. Tygress brings those jobs into one self-hosted gateway, with one set of identity, rate-limit, budget and guardrail policies.

Feature comparison of Tygress with Kong, Apigee, LiteLLM, and Portkey (Palo Alto Networks)
CapabilityTygressKongApigeeLiteLLMPortkey (Palo Alto Networks)
AvailabilityPre-releaseGenerally availableGenerally availableGenerally availableGenerally available
API gatewayPass-through only
AI / LLM gatewayBuilt inSeparate AI Gateway 2.xManagedNativeNative
Gateway-held approvalsMCP + A2A, default denyNot documentedNot documentedClient-sideNot documented
Enterprise identity (OIDC/LDAP/mTLS)OIDC/mTLS: EnterpriseOIDC/mTLS (no LDAP)PartialPartial
MCP gatewayBuilt inAI licenseManagedNative
A2A agent gatewayBuilt inAI licenseNot documentedAgent Gateway
Self-hosted / air-gappedYes (AI 2.x needs Konnect)Hybrid; Private Cloud air-gapYes (air-gap: Enterprise)OSS + hybrid; no air-gap
AI governance & DLPBuilt inAI pluginsManagedBuilt in + integrationsPartial
Multi-provider failoverAI licenseManaged
Semantic cachingAI licenseManagedVia cache backendPaid tiers
Post-quantum TLSHybrid key exchange by default; per-route enforcement; no PQ certificatesHybrid key exchange + ML-DSA certificates (non-FIPS)Client leg via Cloud LBNot documentedNot documented
Data-plane runtimeRust (own transport)NGINX / LuaManaged / hybridPython + Rust (beta)Node/TS
Open sourceCore planned at launchLimitedYesPartial

Tygress column reflects the current pre-release build. Competitor columns were checked against public documentation on 6 October 2026. “AI license” marks Kong features its documentation labels “AI License Required”. Portkey is now part of Palo Alto Networks as Prisma AIRS AI Gateway. Product editions and competitor features change; verify requirements during procurement.

Pricing

Every gateway feature in the $0 tier

Planned pricing. Paid plans add managed hosting, support, SLAs and services, not gateway features.

Open-source core

Planned at launch

Self-hosted

$0

The whole gateway on your own infrastructure: every capability on this page.

Join waitlist
  • HTTP/1.1, HTTP/2, gRPC, gRPC-Web, gRPC-JSON transcoding, WebSocket, TCP/UDP/TLS passthrough
  • Routing, load balancing, active and passive health checks, retries, circuit breaking
  • Rate limiting (token bucket, fixed window), shared across replicas with Redis
  • 60+ built-in plugins, including OIDC, JWT, LDAP/AD, mTLS, OAuth2 introspection, JWE and OPA
  • Sandboxed proxy-wasm plugins and inline Rhai or JavaScript
  • AI gateway: 9 provider types through Chat Completions, Responses and Messages
  • Virtual keys, token limits, USD budgets for keys and teams
  • Prompt guard, DLP redaction, model allowlists, semantic cache
  • MCP gateway, A2A routing, human approvals
  • Prometheus, OpenTelemetry (GenAI), Langfuse, append-only PostgreSQL records of config revisions and approval decisions
  • Hybrid post-quantum TLS key exchange, ACME certificates, zero-downtime upgrades
  • Operator console with OIDC SSO (opt-in build)

Planned services

Cloud

Planned

Managed by Tygress

$49/mo

Tygress runs and upgrades the gateway for you.

  • The full gateway, hosted by Tygress
  • Managed upgrades
  • Uptime SLA
  • Email support
Join waitlist

Cloud (Self-Managed)

Planned

Your infrastructure

$99/mo

Data planes in your cloud account; the control plane managed by Tygress.

  • The full gateway, with data planes in your own cloud account
  • Managed control plane, hosted by Tygress
  • Priority support with a response-time SLA
  • Onboarding help
Join waitlist

Enterprise

Planned

Services and SLAs

Custom

For teams running Tygress on-premises or air-gapped who need a support contract.

  • Dedicated support engineer
  • Custom SLAs
  • Architecture review
  • Migration help from your current gateway
  • Onboarding and training
  • Custom plugin development (proxy-wasm)
Join waitlist

All prices are preliminary and may change before launch. An open-source core is planned at launch. Release and license details will be published with the release.

FAQ

Frequently asked questions

Availability, current capabilities, and how Tygress fits your API and AI workloads.

What is Tygress?

Tygress is a self-hosted API gateway and AI gateway written in Rust on its own purpose-built async transport. It routes REST, gRPC and WebSocket APIs, LLM requests, MCP tools and A2A agents through one data plane with shared identity, rate limiting, cost and guardrail policies, and also proxies TCP/UDP traffic at L4.

Is Tygress available to use today?

Tygress is pre-release. The gateway is built and in active pre-1.0 development, but it is not publicly available yet. Join the early-access waitlist for beta availability and product updates.

Can one gateway handle both API and AI traffic?

Yes, if it understands both. An API gateway manages identity and traffic for REST and gRPC services; an AI gateway adds model routing, provider credentials, token budgets, and AI guardrails. Tygress runs both in one self-hosted data plane, so one identity, rate-limit, and logging model covers APIs, models, and agents. Whether you should consolidate depends on your workloads and existing infrastructure.

Is Tygress built on Pingora?

No. Earlier builds used a vendored fork of Cloudflare's Pingora. Since September 2026, Tygress runs on tygress-proxy, its own async transport, and its own request engine, written in Rust. Owning the transport lets DLP and prompt guards inspect every chunk of a streamed reply, and runs provider translation and request signing inside each retry and failover attempt, after every plugin edit.

Can I self-host Tygress or run it air-gapped?

Tygress runs as one binary, or as separate control-plane and data-plane roles that sync configuration over NATS. The source includes a Dockerfile for a distroless, non-root container image and a Docker Compose reference stack; no image is published yet. An opt-in Kubernetes build watches EndpointSlices to track upstream pods. It is designed to run in a private VPC or on-premises; fully air-gapped operation also requires local model endpoints, identity services, and other dependencies.

Will AI requests stay inside my network?

Self-hosting the gateway gives you control of routing and policy. Requests sent to external model APIs still leave your network; keeping inference local requires local model endpoints and control of logs, caches, and telemetry.

Which LLM providers and APIs does Tygress support?

Tygress connects to OpenAI, Anthropic, Google Gemini, Google Vertex AI (Gemini models), Azure OpenAI, AWS Bedrock, Mistral, Groq, and any OpenAI-compatible endpoint. Clients can call any of them through the OpenAI Chat Completions, OpenAI Responses, or Anthropic Messages API: natively where the provider supports that API, and bridged through chat completions otherwise. Chat Completions reaches Anthropic, Gemini, and Vertex through each vendor's own OpenAI-compatible endpoint. Some provider-specific features, such as stored conversations, hosted tools, prompt caching, or image input, need a native endpoint.

How does Tygress control LLM costs?

Tygress reads token usage from provider replies, including cached and reasoning tokens where the provider reports them, and prices it from a bundled 80-model catalog or your own pricing. Token limits and USD budgets run over fixed, sliding, rolling, or token-bucket windows keyed per consumer, route, or any request variable. Durable daily, monthly, or lifetime caps apply per consumer and per team, with webhook alerts at thresholds you set. Semantic caching can answer a similar non-streaming prompt without calling the chat model, and semantic routing, once switched from shadow mode to enforce, can send simpler prompts to cheaper models. Actual savings depend on your models, traffic, and cache settings.

What security controls does Tygress include?

Authentication plugins cover OIDC, JWT with remote JWKS, LDAP and Active Directory, mTLS, OAuth2 token introspection, JWE, HMAC, basic auth, and API keys, with OPA and ACLs for authorization. AI controls include PII and secret redaction, a prompt-injection guard, model and MCP tool allowlists, and human approvals for MCP tool calls and A2A methods, with each decision recorded in an append-only PostgreSQL table when a database is configured. These controls support your security policies; deployment choices and upstream services also affect how data is handled.

What isn't built yet?

Tygress is pre-1.0, and some things are not done. There is no published container image, Helm chart, or Kubernetes manifest, and the operator console, Kubernetes discovery, Kafka logging, and the local embedder for the semantic cache are opt-in builds. HTTP/3, WebSocket over HTTP/2, SAML, post-quantum certificates and signatures, and Windows are not supported. The OpenAI Realtime API returns 501, failover between A2A agents is not wired, and some proxy-wasm hostcalls (HTTP callouts, shared data and queues, timers) are not implemented. Benchmarks will be published with their methodology. The configuration schema and admin API may change before 1.0.

Will Tygress be open source?

An open-source core is planned at launch. Release and license details will be published with the release.

What can Tygress do?

Current capabilities include: API gateway for HTTP/1.1, HTTP/2, gRPC (including gRPC-Web and JSON transcoding), WebSocket, and TCP/UDP/TLS passthrough, with load balancing, active and passive health checks, retries, and circuit breaking; 60+ built-in plugins for authentication (OIDC, JWT, LDAP, mTLS, OAuth2 introspection), authorization (ACL, OPA), traffic control, and logging, plus sandboxed proxy-wasm plugins and inline Rhai or JavaScript; LLM routing across nine provider types with failover, virtual keys, token limits, USD budgets for keys and teams, and a bundled model-pricing catalog; AI guardrails: prompt-injection detection, DLP that redacts, blocks, or audits PII and secrets in prompts and can redact them in buffered and streamed replies, model allowlists, and semantic caching; MCP tool gating and human approvals, REST-to-MCP toolsets, and A2A agent routing; Prometheus metrics, OpenTelemetry tracing with GenAI semantic conventions, Langfuse export, and append-only PostgreSQL records of config revisions, approval decisions, and autopilot changes; Hybrid post-quantum TLS key exchange (X25519MLKEM768), offered by default and enforceable per listener, upstream, or route; ACME certificates; and zero-downtime binary upgrades.

Related guide: AI gateway vs API gateway

Pre-release

Join the Tygress early-access waitlist

Evaluating a gateway for your APIs or AI workloads? Tygress is pre-release. The gateway is built and in active pre-1.0 development, but it is not publicly available yet. Join the early-access waitlist for beta availability and product updates.

Beta availability updates
No payment required
No spam, ever