Skip to main content
Connect your client
Plans and pricing

Monthly plans for the MCP compression gateway.

A compression is one POST /v1/compress call (or one MCP tool call that wraps it), capped at the per-tier document size below. Five plans: from a free developer tier through reserved-capacity Enterprise. All MCP tools included on every paid plan.

Prices in USD. Annual contracts on Business and above. Procurement artifacts, sub-processor list, and DPA are in Procurement below.

What compression saves in practice
Pro · $49/mo
~$375/mo in token cost
Claude Sonnet 4 at $3/MTok input
at 50% avg reduction, 5K avg doc
Team · $99/mo
~$750/mo in token cost
Claude Sonnet 4 at $3/MTok input
pooled across unlimited seats
Business · $199/mo
~$3,750/mo in token cost
Claude Sonnet 4 at $3/MTok input
metered overage past 500K limit
Enterprise
Custom ROI projection
Any model, any scale
we model your traffic, you see the delta

Estimates use a conservative 50% typical reduction (see live data at /v1/global-savings ) × Claude Sonnet 4 list rate ($3/MTok input) × 5K avg token doc. Your actual savings depend on doc mix and model. Use the calculator below to project your volume.

Tier estimator

What does your usage cost?

Drag the slider to your expected monthly compression volume. We’ll recommend a tier and project your monthly cost. Real reduction depends on document mix; the calculator gives the floor.

Billing cadence
50K
5005M
Recommended tier
Pro
Effective monthly cost
$49/mo
Tier limits
Up to 50K compressions / month · 1 MB max doc size · 30 days payload retention · 1 seat.
Start Pro

Numbers are calculator estimates. Real reduction depends on document mix, fidelity, and downstream model.

Live production data from /v1/global-savingsTypical saving: ~50% token reductionMCP tools: 193 included on every paid planUptime: gotcontext.ai/status
Billing cadence (Pro and Team)
Free: 1 seatPro: 1 seatTeam, Business, Enterprise: unlimited seats, flat price
Free
Solo dev · try before you buy
$0/mo

For individuals validating semantic compression on their own inputs.

Start free

Works with Claude Code, Cursor, and any MCP client

Included
  • 1,000 compressions / month (hard-stop at 1,200)
  • 100 KB max document size
  • 1 concurrent compression slot
  • 23 core MCP tools (compression + search_semantic + gc_pre_flight)
  • 14-day payload retention
  • Community support · no SLA
Most Popular
Pro
Individual developer · solo AI engineering
$49/mo

For individual developers running 50k context compressions per month.

Start 14-day free trial

14 days free, then $49/mo. Cancel anytime: one click, no questions.

Included
  • 50,000 compressions / month (hard-stop at 60,000)
  • 1 MB max document size
  • 2 concurrent compression slots · 60 req/min rate limit
  • All MCP tools (compression, memory, code analysis, multimodal, ACE)
  • 30-day payload retention
  • Email support · 2 business-day first response · no SLA
Business
Growth-stage company · compliance + self-hosted
$199/mo

Shared infrastructure. Unlimited seats. 99.5% SLA. Metered overage past 500k. Includes SSO (SAML 2.0 + OIDC), audit-log export, DPA, and self-hosted Docker.

Get Business: $199/mo
Included

Everything in Pro, plus:

  • 500,000 compressions / month pooled · metered overage $0.50 per 1,000 (auto-billed)
  • 10 MB max document size
  • 8 concurrent compression slots · 500 req/min rate limit
  • Self-hosted Docker: data plane in your VPC (= BYOK answer)
  • 99.5% monthly SLA with 10/25/50% credit schedule
  • Priority email + Slack-connect · 1 BD first response
Custom scope · Talk to sales
Enterprise
Larger teams · custom scope · self-hosted
Custom

Custom terms scoped to your workload. Talk to us about what you actually need.

Talk to sales
Included
  • Everything in Business
  • Custom rate limits and quotas, scoped per contract
  • Configurable payload retention · zero-retention mode in self-hosted
  • Self-hosted deployment with an Ed25519-signed license (air-gapped OK)
  • Uptime and support terms negotiated per contract
  • Data residency on request (US today, EU on roadmap H2 2026)
Team — $99/mo
  • unlimited seats
  • pooled 100k compressions
  • 5 MB docs
Get Team: $99/mo

Plans differ on volume and fidelity, not capability. All 193 MCP tools ship on every paid plan: compression, semantic memory, code analysis, multimodal, and ACE workflows.

View all 193 tools →
How it compares

Why not just run LLMLingua?

The obvious question. LLMLingua is free and open source. The table below covers what you give up when you self-host vs using a managed MCP gateway.

Comparison: gotcontext vs LLMLingua, Langfuse, and per-token APIs (Cohere/Voyage)
DimensiongotcontextLLMLingua (OSS)Langfuse ($0 to $29)Cohere/Voyage Compact
MCP gateway built in✓ Native: Claude Code, Cursor, any MCP clientBuild it yourselfNot a compression toolAPI call, not MCP-native
Compression engineSemantic (ONNX + PageRank). Local, no LLM API call.Prompt-compression (token-level)No compression; tracing onlyEmbedding model reranking
Setup time< 5 min: add MCP server URL to claude_desktop_config.jsonPython env, GPU recommended, write integration ↗ LLMLingua docs~10 min (SDK + API key)~5 min (API key + write call)
Maintenance burdenZero: managed infra, version upgrades automaticModel updates, infra, embedding drift (your ops team)Low (managed SaaS)Low (managed SaaS)
Self-hosted optionBusiness and above: data plane in your VPCAlways self-hosted (that's the product)$0 self-host or $29/mo cloudCloud API only
Pricing modelPer-compression flat (not per-token): predictable at scaleFree (your infra cost)Free tier / $29 teamPer 1M tokens (variable)

LLMLingua and Langfuse are open source projects we respect. This comparison reflects their architectures, not a claim of superiority. Choose what matches your deployment model and team capacity.

Feature comparison by tier

Feature comparison by tier

Limits, support, and security across all five plans.

Limits
Limits feature comparison by tier: Free, Pro, Team, Business, and Enterprise
SpecificationFreeProTeamBusinessEnterprise
Monthly compressions1,00050,000100,000500,000Unlimited (within capacity pool)
Max document size100 KB1 MB5 MB10 MBCustom (negotiable)
Overage policyHard-stopHard-stop at 60KHard-stop at 120KMetered $0.50 / 1KContractual
Batch ingestion—IncludedIncludedIncludedIncluded
Async batch queue——IncludedIncludedIncluded
Compression projects——IncludedIncludedIncluded
Seats, retention, SLA
Seats, retention, SLA feature comparison by tier: Free, Pro, Team, Business, and Enterprise
SpecificationFreeProTeamBusinessEnterprise
Seats11Unlimited (pooled quota)Unlimited (pooled quota)Unlimited
Payload retention14 days30 days90 days1 yearConfigurable + zero-retention mode
Monthly uptime SLA———99.5%Custom terms
Service-credit schedule———10 / 25 / 50%Custom terms
Status page———gotcontext.ai/statusgotcontext.ai/status
Embeddings
Embeddings feature comparison by tier: Free, Pro, Team, Business, and Enterprise
SpecificationFreeProTeamBusinessEnterprise
Standard compressionIncludedIncludedIncludedIncludedIncluded
Accelerated compression (3-5x faster)—IncludedIncludedIncludedIncluded
Custom embedding models————Included
Security & control
Security & control feature comparison by tier: Free, Pro, Team, Business, and Enterprise
SpecificationFreeProTeamBusinessEnterprise
API key managementIncludedIncludedIncludedIncludedIncluded
API rate limit10 req/min60 req/min300 req/min500 req/minCustom
MCP Server tool access23 core compression toolsAll MCP toolsAll MCP toolsAll MCP toolsAll MCP tools
Fidelity ProfilesIncludedIncludedIncludedIncludedIncluded
Prompt Cache Audit—IncludedIncludedIncludedIncluded
Advanced analytics & CSV export——IncludedIncludedIncluded
Teams——IncludedIncludedIncluded
Webhooks—IncludedIncludedIncludedIncluded
Audit-log export (NDJSON/CSV)———IncludedIncluded
SSO via SAML 2.0 + OIDC———IncludedIncluded
Self-hosted Docker (data plane in your VPC)———IncludedIncluded
BYOK (via self-hosted = your VPC)———IncludedIncluded
SCIM provisioning————On roadmap H2 2026
Customer-managed encryption keys (cloud)————On roadmap H2 2026
Data residency (US / EU / APAC)USUSUSUSUS (EU + APAC on roadmap)
DPA, IP indemnity, custom MSA———IncludedIncluded
Support & billing
Support & billing feature comparison by tier: Free, Pro, Team, Business, and Enterprise
SpecificationFreeProTeamBusinessEnterprise
SupportCommunityEmail · 2 BDEmail · 1 BDPriority email + Slack-connectDedicated channel + named CSM
P1 first response———1 business day4 hours
Payment methods—CardCard + ACHCard + ACH + Wire + PO + InvoiceCustom (invoice / PO / ACH / Wire)
Annual discount—20% (2.4 months free)20% (2.4 months free)Annual invoice onlyCustom contract
Refund policy—14-day money-back14-day money-backAnnual prorated within 30 daysPer contract
Platform
Platform feature comparison by tier: Free, Pro, Team, Business, and Enterprise
SpecificationFreeProTeamBusinessEnterprise
Command Palette (Cmd+K)IncludedIncludedIncludedIncludedIncluded
Activity FeedIncludedIncludedIncludedIncludedIncluded
Dark/Light ThemeIncludedIncludedIncludedIncludedIncluded
CSV ExportIncludedIncludedIncludedIncludedIncluded
Queue Monitor (real-time SSE)—IncludedIncludedIncludedIncluded
Webhook Notifications—IncludedIncludedIncludedIncluded
Usage Analytics—IncludedIncludedIncludedIncluded
GitHub Integration——IncludedIncludedIncluded
RBAC Roles——IncludedIncludedIncluded
Shared Projects——IncludedIncludedIncluded
MCP Tool Compression——IncludedIncludedIncluded
SSO / SAML———IncludedIncluded
Audit Trail———IncludedIncluded
Dedicated Support———IncludedIncluded
Custom Integrations———IncludedIncluded
Savings

Project your savings.

The ~50% figure on the landing hero is a conservative typical saving (live data at/v1/global-savingsruns higher), useful as a directional signal, not a projection of your savings. Real reduction depends on your document mix, fidelity choice, and downstream model. Per-model breakdowns (Opus 4.7 vs Gemini Flash vs GPT-5.5) live at /savings-by-model.

Want a number for your own traffic? Use the tier estimator above to project your monthly cost by compression volume, or contact us with 7 days of usage data and we’ll model the monthly delta against your raw token cost across any model.

Compliance and procurement

Built for procurement review

The artifacts a Fortune-500 vendor risk team will ask for, ready before the call.

  • SOC 2

    Type I in progress, target Q3 2026. Not yet certified; stated honestly.

    View page
  • DPA

    Available on request, emailed within one business day; GDPR Art. 28 conformant.

    View page
  • Sub-processors

    Cloudflare · Fly · Supabase · Upstash · Clerk · Polar · Resend · Sentry · PostHog. Full list with 30-day change notice.

    View page
  • Self-hosted Docker

    Business and Enterprise. Data plane in your VPC; control plane SaaS. Operates as the BYOK answer.

  • Audit log export

    NDJSON + CSV; 90-day retention on Business, configurable on Enterprise.

  • Status page

    gotcontext.ai/status shows current operational status. Required reading before signing any SLA tier.

    View list
  • Liability cap

    Negotiable on annual contracts; default capped at 12 months of fees in MSA template.

  • On roadmap (H2 2026)

    SCIM provisioning · cloud BYOK / CMEK · EU + APAC data residency · SOC 2 Type II close.

Frequently asked questions

Answers before the call

Anything not covered here? Use the contact form below.

One compression is one POST /v1/compress request (or one MCP tool call that wraps it), capped at the per-tier document size: 100 KB Free, 1 MB Pro, 5 MB Team, 10 MB Business, custom on Enterprise. A single 30 KB design doc, a 500 KB GitHub diff, and a 2 MB transcript all count as one compression each, regardless of how many tokens are saved.

Contact sales

Enterprise volume and self-hosted

Compliance reviews, custom SLAs, dedicated capacity, on-prem deployments. Tell us about your use case and we’ll respond within one business day.

Who you are

What you need

Plan interest
Primary requirements(select all that apply)

Anything else

Minimum 10 characters. Include team size, compliance constraints, and timeline if known.

Required

© 2026 gotcontext.aiEffective May 2026Security · Changelog · Docs