Skip to main content
Category AI Infrastructure & Model APIs

AI Infrastructure & Model APIs

Updated June 28, 2026: compare OpenRouter, OpenAI API, Claude API, Gemini API, Mistral, Groq, Replicate, Modal, Browserbase, Firecrawl, Composio, LiteLLM, Respan, OpenLIT, Opik, Inspect AI, Guardrails AI, LM Evaluation Harness, OpenAI Evals, LlamaIndex, Haystack, Agno, Instructor, Chainlit, LangSmith, Braintrust, Patronus AI, DeepEval, Traceloop, Arize Phoenix, Ragas, OpenPipe, LangWatch, BAML, DSPy, Portkey, Zep, promptfoo, Mem0, Tavily, Deepgram, vector databases, and governance tradeoffs.

8/10 Strong
Best model router

Free tier (25+ models, 50 req/day) · Pay-as-you-go (5.5% platform fee on 400+ models) · Enterprise custom

Best model router

OpenRouter

Unified LLM API for hundreds of models, with OpenAI-compatible requests, provider routing, fallbacks, app attribution, and per-model token pricing.

Editorial · no paid placements

Quick paths

Buyer path

Source-backed shortlist

Source
Registered source
Freshness
Current
Confidence
High confidence

Best web data API for agents

Firecrawl

Firecrawl is the better shortlist when a product needs search, scrape, crawl, structured extraction, screenshots, and LLM-ready markdown rather than full browser-session ownership.

Plan
Free for testing, then paid credits after endpoint modeling
Confidence
high
Verified
2026-06-28
Source
Registered source
Freshness
Current
Confidence
High confidence
Verified
Read review Build comparison

Best tool and auth layer for agents

Composio

Composio fits AI products that need app toolkits, managed user authentication, session tools, and hosted MCP access without building every connector internally.

Plan
Free for prototypes, $29 plan for production tests
Confidence
high
Verified
2026-06-28
Source
Registered source
Freshness
Current
Confidence
High confidence
Verified
Read review Build comparison

Best open-source LLM gateway

LiteLLM

LiteLLM is the better shortlist when teams need one OpenAI-compatible interface across 100+ providers, with routing, virtual keys, spend tracking, guardrails, MCP, and enterprise gateway controls.

Plan
Open-source proxy first, Enterprise when SSO, audit logs, support, or multi-team governance matter
Confidence
high
Verified
2026-06-28
Source
Registered source
Freshness
Current
Confidence
High confidence
Verified
Read review Build comparison

Best framework for agents over data

LlamaIndex

LlamaIndex fits teams building RAG, context augmentation, workflows, and agents over private or domain-specific data, with a managed LlamaCloud/LlamaParse route when parsing and retrieval should be outsourced.

Plan
MIT framework first, LlamaParse or LlamaCloud after document-processing volume is modeled
Confidence
high
Verified
2026-06-28
Source
Registered source
Freshness
Current
Confidence
High confidence
Verified
Read review Build comparison

Best open-source RAG pipeline framework

Haystack

Haystack is the better shortlist when developers want Apache-2.0 components, pipelines, document stores, agents, tools, and retrieval systems with an optional deepset managed-platform route.

Plan
Free Haystack framework first, deepset AI Platform when managed deployment and support matter
Confidence
high
Verified
2026-06-28
Source
Registered source
Freshness
Current
Confidence
High confidence
Verified
Read review Build comparison

Best LangChain-native observability layer

LangSmith

LangSmith is the better shortlist when LangChain or LangGraph teams need hosted traces, evals, prompt workflows, monitoring, deployment, sandboxes, Fleet, and Engine controls in one platform.

Plan
Developer for solo tracing, Plus for teams, Enterprise for self-hosted or hybrid needs
Confidence
high
Verified
2026-06-28
Source
Registered source
Freshness
Current
Confidence
High confidence
Verified
Read review Build comparison

Best evals-first release control

Braintrust

Braintrust is the better shortlist when teams need datasets, experiments, traces, scores, prompt testing, monitoring, and human review tied to release decisions.

Plan
Starter for prototypes, Pro for shared AI teams
Confidence
high
Verified
2026-06-28
Source
Registered source
Freshness
Current
Confidence
High confidence
Verified
Read review Build comparison

Best open-source AI observability platform

Arize Phoenix

Arize Phoenix is the better shortlist when OpenTelemetry traces, evals, prompt iteration, datasets, and experiments need to sit close to engineering workflows.

Plan
Phoenix self-hosted for validation, AX Pro for hosted small-team usage
Confidence
high
Verified
2026-06-28
Source
Registered source
Freshness
Current
Confidence
High confidence
Verified
Read review Build comparison

Best OpenTelemetry-first LLM observability stack

OpenLIT

OpenLIT is the better shortlist when engineers want Apache-2.0 self-hosted traces, metrics, token and cost tracking, prompt workflows, evals, dashboards, GPU monitoring, and OpenTelemetry alignment.

Plan
Free self-hosted edition until managed cloud pricing is public
Confidence
high
Verified
2026-06-28
Source
Registered source
Freshness
Current
Confidence
High confidence
Verified
Read review Build comparison

Best OSS-to-hosted agent eval platform

Opik

Opik fits teams that want open-source or hosted agent traces, Test Suites, assertions, LLM-as-judge metrics, annotation, production monitoring, and Comet-hosted span retention.

Plan
Opik OSS or Free Cloud first, Pro Cloud at $19/month when spans or team needs grow
Confidence
high
Verified
2026-06-28
Source
Registered source
Freshness
Current
Confidence
High confidence
Verified
Read review Build comparison

Best custom agent and safety eval framework

Inspect AI

Inspect AI fits teams that need code-defined evaluations for coding, agentic tasks, reasoning, knowledge, behavior, multimodal understanding, tools, scorers, and sandboxed runs.

Plan
Free MIT framework, with model calls, compute, sandboxes, storage, and review time separate
Confidence
high
Verified
2026-06-28
Source
Registered source
Freshness
Current
Confidence
High confidence
Verified
Read review Build comparison

Best runtime validation guardrails

Guardrails AI

Guardrails AI fits teams that need reusable validators, input and output guards, Pydantic-style structured data, on-fail policies, and Guardrails Hub installs around LLM apps.

Plan
Free Apache-2.0 framework, with hosted or remote-validator pricing confirmed directly
Confidence
high
Verified
2026-06-28
Source
Registered source
Freshness
Current
Confidence
High confidence
Verified
Read review Build comparison

Best standardized benchmark harness

LM Evaluation Harness

LM Evaluation Harness fits model teams that need reproducible academic and leaderboard-style benchmark runs across local models, hosted APIs, vLLM, SGLang, Hugging Face, and other backends.

Plan
Free MIT framework, with GPU, API, backend, storage, and review costs separate
Confidence
high
Verified
2026-06-28
Source
Registered source
Freshness
Current
Confidence
High confidence
Verified
Read review Build comparison

Best code-first RAG evaluation framework

Ragas

Ragas fits developer teams that want open-source metrics, synthetic test data, experiments, and cost-aware eval loops in code rather than a hosted dashboard first.

Plan
Free Apache-2.0 framework, with evaluator model costs modeled separately
Confidence
high
Verified
2026-06-28
Evidence Ragas docs
Source
Registered source
Freshness
Current
Confidence
High confidence
Verified
Read review Build comparison

Best fine-tuning workflow for cheaper specialized models

OpenPipe

OpenPipe is the better shortlist when production request logs can become datasets, fine-tunes, DPO runs, evaluations, and hosted inference for cost or latency reduction.

Plan
Usage-based hosted inference after a baseline eval
Confidence
high
Verified
2026-06-28
Source
Registered source
Freshness
Current
Confidence
High confidence
Verified
Read review Build comparison

Best open-source LLMOps platform

LangWatch

LangWatch fits teams that want traces, evaluations, datasets, AI gateway workflows, DSPy optimization, self-hosting, and monitoring in one open-source LLMOps surface.

Plan
Developer for early checks, Growth for paid capacity, Enterprise for governed rollout
Confidence
high
Verified
2026-06-28
Source
Registered source
Freshness
Current
Confidence
High confidence
Verified
Read review Build comparison

Best combined LLMOps and gateway challenger

Respan

Respan is the better shortlist when a team wants traces, metrics, evals, prompt management, monitors, spend limits, and an OpenAI-compatible gateway tied to the same production traffic.

Plan
Free for instrumentation checks, Team after shared workflow fit is proven, Enterprise for compliance or custom support
Confidence
high
Verified
2026-06-28
Source
Registered source
Freshness
Current
Confidence
High confidence
Verified
Read review Build comparison

Best enterprise AI reliability lab

Patronus AI

Patronus AI fits teams that need managed evals, traces, datasets, guardrails, Percival-assisted eval creation, and current Digital World Model research around agent simulation.

Plan
Developer for testing, Enterprise for production reliability controls
Confidence
high
Verified
2026-06-28
Source
Registered source
Freshness
Current
Confidence
High confidence
Verified
Read review Build comparison

Best open-source LLM evaluation framework

DeepEval

DeepEval is the better shortlist when developer-owned LLM, RAG, agent, chatbot, image, safety, and CI evals should live in code before a hosted quality platform is added.

Plan
DeepEval open source first, Confident AI when hosted team workflows matter
Confidence
high
Verified
2026-06-28
Source
Registered source
Freshness
Current
Confidence
High confidence
Verified
Read review Build comparison

Best OpenTelemetry LLM observability layer

Traceloop

Traceloop fits teams that want OpenLLMetry, OpenTelemetry traces, quality checks, prompt changes, real-time alerts, and a ServiceNow AI Control Tower path around agent runtime evidence.

Plan
Free Forever for 50K-span testing, Enterprise for production scale
Confidence
high
Verified
2026-06-28
Source
Registered source
Freshness
Current
Confidence
High confidence
Verified
Read review Build comparison

Best LLM gateway governance layer

Portkey

Portkey fits production AI teams that need routing, observability, prompt management, guardrails, API key control, budgets, caching, and enterprise gateway policy.

Plan
Production for live apps, Enterprise for private hosting and compliance
Confidence
high
Verified
2026-06-28
Source
Registered source
Freshness
Current
Confidence
High confidence
Verified
Read review Build comparison

Best red-team and eval harness

promptfoo

promptfoo is the sharper first stop when the infrastructure risk is jailbreak testing, vulnerability scanning, guardrails, model security, MCP exposure, or repeatable local evals.

Plan
Community for local tests, Enterprise for shared security workflows
Confidence
high
Verified
2026-06-28
Source
Registered source
Freshness
Current
Confidence
High confidence
Verified
Read review Build comparison

All tools in AI Infrastructure & Model APIs

  1. 1
    Hugging Face Open AI collaboration hub for models, datasets, Spaces, inference endpoints, evaluations, and enterprise ML workflows.
    Free hub access; Pro $9/mo; Team $20/user/mo; Enterprise from $50/user/mo; paid compute/storage 9.3/10
    Try Hugging Face free
  2. 2
    LiteLLM Open-source LLM gateway and Python SDK for one OpenAI-compatible interface across 100+ model providers, with routing, virtual keys, spend tracking, guardrails, MCP, and enterprise controls.
    Free MIT core outside enterprise directory; Enterprise custom 8.8/10
    See LiteLLM pricing
  3. 3
    promptfoo Open-source LLM evaluation, red teaming, vulnerability scanning, guardrails, model security, MCP proxy, code scanning, and enterprise AI security testing.
    Community free / Enterprise custom / On-Premise custom 8.8/10
    Try promptfoo free
  4. 4
    LlamaIndex Open-source framework and managed LlamaCloud stack for building LLM agents over private data, RAG, document parsing, extraction, indexing, retrieval, workflows, and context augmentation.
    Framework free MIT / LlamaParse Free 10K credits / Starter $50/month / Pro $500/month / Enterprise custom 8.5/10
    Try LlamaIndex free
  5. 5
    Arize Phoenix Open-source AI observability, tracing, evaluation, prompt engineering, experiments, and Arize AX hosting for teams improving LLM systems.
    Phoenix open source / AX Pro $50/month / AX Enterprise custom 8.3/10
  6. 6
    Braintrust AI evaluation, tracing, prompt playground, datasets, experiments, monitoring, human review, and observability infrastructure for teams shipping LLM products.
    $0 Starter / $249 Pro / Enterprise custom, plus usage meters 8.3/10
    Try Braintrust free
  7. 7
    Inspect AI MIT-licensed evaluation framework from the UK AI Security Institute and Meridian Labs for coding, agent, reasoning, knowledge, behavior, and multimodal model evals.
    Free MIT framework; model APIs, compute, sandboxes, and storage billed separately 8.3/10
    Try Inspect AI
  8. 8
    LangSmith LangChain's hosted agent and LLM observability platform for tracing, monitoring, evaluation, prompt workflows, deployment, sandboxes, Fleet agents, and Engine optimization.
    $0 Developer / $39 Plus / Enterprise custom, plus usage meters 8.3/10
    Try LangSmith free
  9. 9
    LM Evaluation Harness MIT-licensed EleutherAI framework for running standardized language-model benchmarks across local models, APIs, vLLM, SGLang, Hugging Face, and leaderboard tasks.
    Free MIT framework; model APIs, GPUs, backend extras, and storage billed separately 8.3/10
  10. 10
    Modal Serverless cloud for Python, GPUs, jobs, web endpoints, sandboxes, queues, and AI apps that should scale without managing infrastructure.
    Starter $0 with $30/mo credits; Team $250/mo plus compute; GPU billed per second 8.3/10
    Try Modal free
  11. 11
    Portkey LLM gateway, observability, prompt management, routing, guardrails, governance, caching, and cost controls for production AI applications.
    Developer free / Production $49/month / Enterprise custom 8.3/10
    Try Portkey free
  12. 12
    Together AI AI infrastructure platform for serverless inference, dedicated GPU deployments, fine-tuning, code sandboxes, and open-model training workflows.
    Serverless tokens; dedicated H100 $6.49/hr, H200 contact sales, B200 $11.95/hr; GPU clusters H100 $5.49/hr, H200 $6.79/hr, B200 $9.95/hr; sandbox $0.03/session 8.3/10
    Try Together AI
  13. 13
    Weaviate Open-source vector database and managed cloud for RAG, semantic search, hybrid search, multi-tenancy, embeddings, and AI-native retrieval.
    Free self-host/cloud entry; Flex from $45/mo; Plus from $280/mo; Premium from $400/mo; AI services usage-based 8.3/10
    Try Weaviate
  14. 14
    DeepEval Open-source LLM evaluation framework from Confident AI for metrics, test cases, RAG evals, agent evals, tracing, datasets, and CI-friendly quality gates.
    DeepEval open source / Confident AI Free / Starter $9.99/user/month / Team and Enterprise custom 8/10
    Try DeepEval free
  15. 15
    Haystack Apache-2.0 AI orchestration framework from deepset for production LLM apps, RAG systems, agents, multimodal search, reusable components, pipelines, tools, and document stores.
    Haystack framework free Apache-2.0 / deepset AI Platform starts free / Enterprise custom 8/10
    Try Haystack free
  16. 16
    OpenRouter Unified LLM API for hundreds of models, with OpenAI-compatible requests, provider routing, fallbacks, app attribution, and per-model token pricing.
    Free tier (25+ models, 50 req/day) · Pay-as-you-go (5.5% platform fee on 400+ models) · Enterprise custom 8/10
    Try OpenRouter free
  17. 17
    Pinecone Managed vector database for semantic search, hybrid search, RAG, recommendations, Pinecone Assistant, and production AI retrieval workloads.
    Free Starter, $20/mo Builder, $50/mo Standard minimum, $500/mo Enterprise minimum plus usage 8/10
    Try Pinecone free
  18. 18
    Qdrant Open-source vector database written in Rust, with managed cloud, Free/Standard/Premium tiers, hybrid/private cloud options, metadata filtering, payload indexes, and RAG-ready retrieval.
    Free self-host; Free Cloud tier; Standard usage-based; Premium/Hybrid/Private sales-led 8/10
    Try Qdrant
  19. 19
    Ragas Open-source evaluation framework for LLM apps, RAG systems, metrics, synthetic test data, experiments, and cost-aware eval loops.
    Free open-source framework; model/evaluator usage costs vary 8/10
    Try Ragas
  20. 20
    Replicate Developer platform for running open and hosted AI models by API, with official models, community models, custom deployments, and usage-based pricing.
    Usage-based by official model output or hardware runtime 8/10
    Try Replicate
  21. 21
    LangWatch Open-source LLMOps platform for traces, evaluations, datasets, AI gateway workflows, DSPy optimization, self-hosting, and monitoring.
    Developer free / paid plans from €59/month / Enterprise custom 7.8/10
    Try LangWatch free
  22. 22
    Mem0 Memory layer for AI agents that persists user, session, and agent context across conversations, with a managed Platform and Apache-2.0 open-source self-hosting path.
    Free 10K memories, Starter $19/mo, Growth $79/mo, Pro around $249-$250/mo, Enterprise custom 7.8/10
    Try Mem0 free
  23. 23
    OpenLIT Apache-2.0, OpenTelemetry-native LLM observability platform for traces, metrics, costs, prompts, evals, dashboards, and GPU monitoring.
    Self-hosted $0 forever; managed cloud coming soon with no public price verified 7.8/10
    Try OpenLIT
  24. 24
    OpenPipe Fine-tuning, request logging, datasets, evaluations, DPO, and hosted inference for turning expensive prompts into cheaper specialized models.
    Usage-based hosted inference from $0.48 per 1M tokens for <=8B models; Enterprise custom 7.8/10
    See OpenPipe pricing
  25. 25
    Opik Open-source and hosted AI observability and evaluation platform from Comet for agent traces, test suites, LLM-as-judge metrics, and production monitoring.
    OSS $0; Free Cloud $0; Pro Cloud $19/month; Enterprise custom 7.8/10
    Try Opik free
  26. 26
    Patronus AI AI evaluation and simulation infrastructure for LLM apps, agent debugging, evaluators, traces, datasets, prompts, guardrails, and Digital World Models.
    Developer free with $10 API credits / evaluator API calls from $10 per 1K / Enterprise custom 7.8/10
    Try Patronus AI free
  27. 27
    Traceloop OpenTelemetry-based LLM observability and evaluation platform built on OpenLLMetry for traces, quality checks, prompt management, experiments, and enterprise AI monitoring.
    Free Forever $0/month up to 50K spans / Enterprise custom 7.8/10
    Try Traceloop free
  28. 28
    Zep Production agent memory and context engineering platform with temporal context graphs, credit-based plans, hosted cloud, BYOK, BYOC, and Graphiti open-source context graph work.
    Free 10k credits/month / Flex $125/month / Flex Plus $375/month / Enterprise custom 7.8/10
    Try Zep free
  29. 29
    Composio Tool-calling and MCP infrastructure for AI agents, with 1000+ app toolkits, managed authentication, session tools, hosted MCP URLs, and usage-based pricing.
    $0-$229/month plus Enterprise 7.5/10
    Try Composio free
  30. 30
    Guardrails AI Apache-2.0 guardrails framework and Hub for validating, structuring, and quality-controlling LLM inputs and outputs with reusable validators.
    Free Apache-2.0 framework; hosted or remote-validator pricing not publicly verified 7.5/10
    Try Guardrails AI
  31. 31
    Firecrawl Web data API for AI agents that search, scrape, crawl, parse, monitor, and interact with pages, then return LLM-ready markdown, HTML, screenshots, or structured data.
    Free tier, paid credit plans, and Enterprise 7.3/10
    Try Firecrawl free
  32. 32
    Respan LLM engineering platform, formerly Keywords AI, for observability, evals, prompt management, and an OpenAI-compatible gateway across many model providers.
    $0 Free; Team paid with additional-seat pricing shown; Enterprise custom 7.3/10
    Try Respan free
  33. 33
    OpenAI Evals OpenAI evaluation framework and platform for testing prompts, models, graders, and LLM app regressions, with a scheduled platform shutdown on November 30, 2026.
    MIT open-source repo; OpenAI API usage billed separately; platform Evals scheduled to shut down November 30, 2026 5/10
    Try OpenAI Evals
  34. 34
    DSPy MIT-licensed framework from Stanford for programming, optimizing, and evaluating language-model systems with signatures, modules, metrics, optimizers, agents, and structured inputs/outputs.
    Free MIT framework; model/provider, data, eval, and hosting costs separate 8.3/10
    Try DSPy
  35. 35
    BAML Apache-2.0 language and toolchain from BoundaryML for typed LLM functions, generated clients, structured outputs, robust parsing, tests, streaming, multimodal inputs, and Boundary Studio traces.
    Free Apache-2.0 framework; model/provider and optional Boundary Studio costs separate 8/10
    Try BAML
  36. 36
    Browserbase Cloud browser infrastructure for agents, scraping, QA automation, and web data workflows that need managed Chromium, Fetch, Search, identity, runtime, and observability.
    $0, $20/mo, $99/mo, or custom scale plans plus usage 8/10
  37. 37
    Instructor MIT-licensed structured-output library for getting validated JSON from LLMs with Pydantic-style models, retries, and provider adapters.
    Free MIT library; model/provider, hosting, and validation costs separate 8/10
    Try Instructor
  38. 38
    Pydantic AI MIT-licensed Python agent framework from the Pydantic team, built around typed agents, structured outputs, tools, dependencies, MCP, evals, graph workflows, and Logfire observability.
    Free MIT-licensed framework; model, infrastructure, and optional Logfire costs separate 8/10
    Try Pydantic AI
  39. 39
    Agno Apache-2.0 agent platform SDK and AgentOS control plane for building, running, observing, and managing production agent systems in your own stack.
    Free open source; Pro $150/month; Enterprise custom 7.8/10
    See Agno pricing
  40. 40
    Mirascope MIT-licensed provider-agnostic SDKs for typed LLM calls, tools, structured outputs, streaming, agents, and OpenTelemetry-friendly observability.
    Free MIT SDKs; Mirascope Cloud discontinued 7.8/10
    Try Mirascope
  41. 41
    Outlines Apache-2.0 structured-generation library from .txt for constraining LLM output to JSON Schema, regex, grammars, and typed schemas during generation.
    Free Apache-2.0 library; Dottxt API access available by request with no public rates verified 7.8/10
    Try Outlines
  42. 42
    Tavily Real-time search, extract, crawl, map, and research APIs for AI agents and RAG workflows, priced by API credits.
    Free 1,000 credits/month, $0.008/credit PAYG, $30-$500 monthly plans, Enterprise custom 7.8/10
    Try Tavily free
  43. 43
    Dify Open-source platform for building AI apps, agents, chatbots, workflows, and RAG systems, with Dify Cloud plus self-hosted Community, Premium, and Enterprise routes.
    Free Sandbox, paid cloud plans, and self-hosted editions 7.3/10
    Try Dify free
  44. 44
    Flowise Open-source visual builder for AI agents and LLM workflows, with chatflows, agentflows, assistants, RAG pipelines, evaluations, tracing, teams, and self-hosting.
    Open-source self-hosting; Flowise Cloud Free and Starter routes 7.3/10
    Try Flowise
  45. 45
    Chainlit Apache-2.0 Python framework for building conversational AI interfaces, prototypes, and internal chat apps around LLM workflows.
    Free Apache-2.0 framework; model, hosting, and support costs separate 6.8/10
    Try Chainlit

Quick Decision

AI infrastructure tools sit underneath the apps people see. They route model calls, host open models, run GPU workloads, store embeddings, power RAG, transcribe audio, generate media, and help teams compare cost, latency, quality, and control without rebuilding the stack every month.

This category is for developer and platform buyers. If the user is choosing a chatbot, start with AI Chatbots. If the team is shipping an AI product, agent, retrieval layer, or model-backed workflow, this is the better lane.

The late-May infrastructure update is agent control. CoreWeave’s training-to-inference loop pushes traces, evals, RL, inference, and W&B tooling into one reliability story. OpenAI’s Rosalind Biodefense trusted-access expansion shows that specialist frontier models may ship as gated capability programs. Sysdig’s LLM-agent intrusion report makes runtime telemetry and least-privilege design part of infrastructure buying, not only security cleanup.

The June 3 update widens that control story. Microsoft Build put Work IQ and Foundry around enterprise agents; GitHub made the Copilot SDK generally available while AI Credits became the agent-usage meter; NVIDIA pushed enterprise agents, Cosmos 3, open physical-AI agent skills, Alpamayo 2 Super, RTX Spark, and DGX Station for Windows; Postman launched AI Engineer for API work; RelationalAI moved agentic decision intelligence deeper into Snowflake; 7AI kept security agents in the proactive-hunting lane; and the White House AI cybersecurity order put advanced AI cyber capability into public-sector and critical-infrastructure policy. Infrastructure buyers should evaluate agent stacks by context access, runtime isolation, traces, evals, spend controls, simulation/data pipelines, local-vs-cloud compute, and write-action approvals.

The June 24 update keeps model availability as a first-class infrastructure risk. Claude Fable/Mythos access is still not a normal public self-serve buyer route, GPT-5.2 is retired from ChatGPT, and OpenAI faces reported state-AG scrutiny. Direct frontier API buyers should now document the exact model route in production, the fallback route if a model is suspended, account-specific, preview-only, or retired, the retention policy for that model class, the staff/client access exposure for restricted routes, and the legal/privacy review path for sensitive users. The AI Model Availability & Churn Tracker is now the canonical AiPedia surface for these app/API/router distinctions.

The June 16 infrastructure update is governed data agents. Google Cloud’s data-agent rollout puts Conversational Analytics, Data Engineering Agent, Looker agents, Gemini Enterprise data access, Data Agent Kit, Managed MCP Servers for Databases are GA, while many of the more ambitious analytics, Looker, Gemini Enterprise, and commerce routes remain preview. Infrastructure teams should evaluate these by IAM scope, roles/mcp.toolUser, service permissions, separate production identities, SQL verification, BigQuery spend limits, job labels, audit logging, Model Armor payload logging, and GA-versus-preview fallback plans.

Use OpenRouter when you need one API across many model providers. The current pricing page lists pay-as-you-go access to 400+ models and 60+ providers, with budget controls, activity logs, prompt caching, preferred vendor selections, and model-priced token billing. Its May 27 funding signal makes the category clearer: routing, fallback, governance, and spend visibility are becoming production infrastructure, not just developer convenience.

Use direct vendor APIs when native features matter. OpenAI API is the default direct route for broad multimodal app work. Claude API is the direct route for long reasoning, writing, code, and document workflows. Gemini API inputs, or Veo video generation are part of the product. The June 22 Gemini recheck keeps Gemini 3.5 Flash pricing mode-specific: standard, batch/flex, priority, grounding, tools, and media rows need separate cost modeling.

Use Mistral AI or Groq when price/performance, open-model strategy, European infrastructure, or low-latency inference matters. The June 24 Groq check adds Qwen 3.6 27B, Kimi buyer math. The June 24 Mistral check keeps the timeline and cost model honest: Mistral 3 officially launched on December 2, 2025, while Medium 3.5’s model-card date is April 28, 2026. Current Mistral pricing lists Large 3 at $0.50/M input and $1.50/M output, Medium 3.5 at $1.50/M and $7.50/M, and Small 4 at $0.10/M and $0.30/M, but the Small 4 model card still lists $0.15/M and $0.60/M, and the pricing FAQ still uses a generic Mistral Large $2/$6 example. Benchmark real prompts, confirm the live Studio quote, and pin exact model IDs before switching because model quality, output length, retries, aliases, and source drift change the bill.

Use Replicate or fal.ai when the job is hosted image, video, audio, 3D, or custom-model inference. The June 9 Replicate check keeps it strongest as a broad model catalog and custom-model deployment layer: public models may bill by hardware time or by input/output, while most private deployments bill setup, idle, and active time unless they are labeled fast-booting fine-tunes. fal is stronger when successful-output billing and fast media APIs are the buyer problem; the June 2 check keeps prepaid credits, queue behavior, failed-output billing, and the 50% batch discount as the key pricing details to model.

Use Fireworks AI when the workload is production inference over open or commercial models., cached-token discounts, batch jobs, dedicated GPU deployments, fine-tuning, and B200/B300 capacity are the actual purchase.

Use Browserbase when the infrastructure problem is web interaction.. It belongs here when agents need reliable browser sessions, Fetch/Extract, replay, and model routing rather than just another LLM.

Use Firecrawl when the infrastructure problem is web data for agents. The June 28 check adds it as the cleaned-content counterpart to Browserbase: search, scrape, crawl, structured extraction, screenshots, Interact, and LLM-ready markdown in one credit-based API. Model the exact endpoint mix before buying because crawls, screenshots, browser minutes, and retries can change the bill quickly.

Use Composio when agents need app actions and user-scoped auth. The June 28 check keeps it in the tool-calling infrastructure lane: 1000+ app toolkits, managed authentication, native session tools, hosted MCP URLs, and usage-based tool-call pricing. It is a developer tool layer, not a no-code workflow canvas.

Use LiteLLM when LLM traffic needs an open-source gateway. The June 28 check adds LiteLLM/proxy lane for one OpenAI-compatible interface across 100+ providers, with routing, virtual keys, spend tracking, budgets, guardrails, and Enterprise controls such as SSO, logs, fallback quality, and enterprise-directory licensing still need review.

Use LangSmith when agent observability, evals, and deployment control need a LangChain-native home. The June 28 check keeps Developer at $0/seat/month, Plus at $39/seat/month, and Enterprise custom, but the buyer risk is broader than seats: traces, extended retention, deployment runs, deployment uptime, Fleet, Engine, sandboxes, and outside model/API spend all need usage limits.

Use Braintrust when the infrastructure problem is eval discipline. Braintrust belongs here when datasets, experiments, prompt comparisons, traces, scores, monitoring, and human review need to decide whether a model, prompt, retrieval, or agent change is safe to ship. Starter is useful for prototypes; Pro at $249/month is the first serious shared-team tier, but usage meters still matter.

Use Arize Phoenix when the infrastructure problem is OpenTelemetry-native AI observability. Phoenix belongs here for teams that need traces, evals, prompt iteration, datasets, and experiments tied to LLM app behavior. Self-hosted Phoenix is the open-source path; AX Pro at $50/month is the hosted small-team route, but span volume and Elastic License 2.0 limits need review.

Use OpenLIT when OpenTelemetry alignment is the first requirement. The June 28 check adds OpenLIT traces, metrics, logs, token and cost tracking, prompt workflows, evals, dashboards, vector database monitoring, and GPU monitoring. The self-hosted product is listed at $0 forever, while OpenLIT Cloud is coming soon with no public price verified.

Use Opik when agent traces and evals need an OSS-to-hosted path. Opik belongs here when teams want trace debugging, Test Suites, assertions, LLM-as-judge metrics, production monitoring, annotation queues, and Comet-hosted retention. Start with OSS or Free Cloud, then model Pro Cloud at $19/month plus span and retention meters before rollout.

Use Inspect AI when evals need to be code-defined and sandboxed. The June 28 check adds Inspect AI as the MIT-licensed UK AI Security Institute and Meridian Labs framework for coding, agentic, reasoning, knowledge, behavior, and multimodal evals. It supports datasets, agents, tools, scorers, Inspect View, a VS Code extension, external agents, MCP or custom tools, and sandbox routes. It is free software, but model calls, compute, sandboxes, storage, and review time still need budgets.

Use Guardrails AI when LLM outputs need runtime validation before downstream systems trust them. Guardrails adds reusable validators, Guards, input/output checks, Pydantic-style structured data, on-fail policies, and Guardrails Hub installs around app code. The Apache-2.0 framework is the validated route; public hosted or remote-validator pricing was not verified in the June 28 check.

Use LM Evaluation Harness when benchmark comparability matters. EleutherAI’s MIT framework is the standard-benchmark lane for 60+ academic benchmarks, hundreds of task variants, custom prompts and metrics, and backend routes across Hugging Face, vLLM, API models, OpenAI-compatible local servers, SGLang, OpenVINO, NeMo, Megatron-LM, and more. Pair it with private product evals before making release decisions.

Use OpenAI Evals only as a migration bridge. OpenAI’s deprecations page says Evals platform deprecation was announced June 3, 2026, existing evals become read-only on October 31, 2026, and the dashboard/API are scheduled to shut down on November 30, 2026. Existing users should inventory, export, or migrate eval coverage; new teams should choose a maintained eval stack instead.

Use LangWatch when the team wants open-source LLMOps with traces and evals. LangWatch sits between local eval frameworks and managed observability products: traces, evaluations, datasets, AI gateway workflows, DSPy optimization, self-hosting, and event-sourcing operations. Paid plans start from EUR59/month, but event and retention meters decide real cost.

Use Respan when LLMOps and gateway routing should share one traffic history. The June 28 check adds Respan as the active Keywords AI successor surface: traces, metrics, evals, prompt management, monitors, spend limits, and an OpenAI-compatible gateway sit in one product lane. Free is the first instrumentation route, Team is the shared paid route, and Enterprise is custom. Confirm the live Team base price, retention, data-handling terms, gateway latency, and provider bills before standardizing.

Use Patronus AI when AI reliability needs managed evals plus frontier simulation research. Patronus covers evaluators, experiments, datasets, comparisons, traces, prompts, annotations, guardrails, and Percival-assisted eval creation, while its current company positioning also highlights Digital World Models for simulating agent actions in digital workflows. Confirm whether procurement is buying eval operations, Digital World Model access, or services.

Use DeepEval when LLM evaluation should live in code. DeepEval evals, agent evals, chatbot tests, safety checks, tracing, and CI gates. Confident AI is the hosted platform route when teams need shared evaluation, observability, red teaming, governance, and security controls.

Use Traceloop when OpenTelemetry alignment is the observability requirement. Traceloop and OpenLLMetry instrument LLM apps with OpenTelemetry traces, quality checks, alerts, prompt and model change testing, and IDE reruns. The ServiceNow acquisition path can help enterprise buyers, but roadmap and procurement details need live confirmation.

Use Ragas when RAG and LLM evaluation should live in code. Ragas is the open-source framework lane for metrics, synthetic test data, experiments, and cost-aware evaluation loops. It is free as a framework, but evaluator model calls, embeddings, generated test data, and human review still need a budget.

Use OpenPipe when fine-tuning can lower cost or latency. OpenPipe belongs here when production logs can become datasets, fine-tuned models, DPO runs, evaluations, and hosted inference. Start only after the team has enough clean examples and baseline evals to prove a specialized model beats prompt-only changes.

Use Portkey when LLM traffic needs a gateway control plane. Portkey, governance, caching, and analytics. The Developer plan is explicitly not the production path; Production and Enterprise should be modeled around recorded logs, retention, overages, private hosting, compliance, and provider spend.

Use Zep when the infrastructure problem is production agent memory. Zep is the context-engineering lane for temporal context graphs, persistent user/session memory, and managed deployment options. It belongs beside Mem0 and vector databases, but buyers should decide whether they need memory extraction and graph context or only document retrieval.

Use promptfoo when AI security testing is part of the infrastructure gate. promptfoo is now positioned around open-source evals, red teaming, guardrails, model security, MCP proxy checks, code scanning, and enterprise/on-premise testing workflows. OpenAI announced a March 2026 acquisition agreement, and promptfoo says it is now part of OpenAI, so procurement should verify contract and data-handling paths.

Use Mem0 when the infrastructure problem is agent memory. Mem0 belongs beside vector databases and agent runtimes when an app needs persistent, scoped memories across users, sessions, and agents. The managed Platform is the faster route; the Apache-2.0 open-source route shifts vector database, model, embedder, hosting, deletion, and privacy controls back to the team.

Use Tavily when agents need real-time web search and extraction as an API. Tavily is closer to a search, extract, crawl, map, and research primitive than a browser infrastructure layer. Model basic versus advanced search, extract batches, mapping, crawl size, research mode, retries, and agent-loop behavior before production.

Use LlamaIndex when agents need structured context over private data. The June 28 check adds it as the framework lane for RAG, context augmentation, data connectors, indexes, retrieval, workflows, and agents over domain-specific data. The framework is MIT-licensed, but LlamaParse/LlamaCloud credits, model calls, embeddings, vector storage, ingestion jobs, and evals still need separate budgets.

Use Haystack when RAG and agent apps should be built as open-source pipelines. Haystack belongs here for Apache-2.0 components, pipelines, document stores, tools, agents, multimodal search, and reusable LLM app architecture. deepset AI Platform is the managed route when deployment, governance, testing, support, security, or dedicated resources matter.

Use DSPy when prompt behavior needs measurable optimization. DSPy is a coding-first framework, but it belongs in infrastructure shortlists when an LLM system needs signatures, modules, metrics, optimizers, and examples instead of hand-tuned prompt sprawl. Budget for model calls and data work before assuming optimization will lower costs.

Use Agno when agents need an owned AgentOS-style runtime. Agno is primarily an automation and agent-platform framework, but it belongs in infrastructure shortlists when teams need agents, teams, workflows, memory, knowledge, storage, traces, audit logs, interfaces, and control-plane management in their own stack. Free open source is the first route; Pro is $150/month for a live AgentOS control plane; Enterprise is custom.

Use Instructor when structured LLM outputs are the infrastructure boundary. Instructor is the MIT-licensed library lane for validated JSON, Pydantic-style schemas, retries, and provider adapters. It is useful when extraction, classification, enrichment, or tool arguments need typed outputs, but it still needs evals, retry budgets, and monitoring.

Use Chainlit when a Python LLM workflow needs a quick conversational interface. Chainlit is an Apache-2.0 framework for prototypes, internal chat tools, and demos around RAG or agent workflows. Treat it as a UI framework, not a managed chatbot platform, and verify maintainer/support route, auth, persistence, hosting, and observability before production use.

Use Deepgram when speech is infrastructure. Deepgram is a better fit for product teams adding STT, TTS, audio intelligence, or voice agents than for creators who only need a one-off transcript.

Use Hugging Face when model discovery, model cards, datasets, Spaces, and managed endpoints need to live in one open-AI collaboration surface. The June 23 pricing check keeps Pro at $9/month, Team at $20/user/month, Enterprise from $50/user/month, storage at $12/TB public and $18/TB private before volume discounts, ZeroGPU on RTX Pro 6000 Blackwell, and Inference Endpoints starting around low hourly CPU pricing.

Buyer Paths

Buyer jobStart withWhyWatch out
Multi-model LLM routingOpenRouterOne API, many providers, spend controls, logs, routingRouter fees and provider policy choices still need governance
Direct frontier LLM APIOpenAI, Claude, or GeminiBest when native model features, support, and procurement matterModel access, retirements, legal/data governance, long context, outputs, tools, and video can change cost and risk quickly
Budget/open-model APIMistral AI or GroqUseful for cost-sensitive, latency-sensitive, and sovereignty-sensitive workloadsRequires benchmarking against your actual prompts, exact model IDs, and current model-card/pricing-page drift
Hosted model catalogReplicatePublic, proprietary, and custom models without owning GPUsHardware-time, output-priced media, and private-model idle billing need separate cost modeling
Fast media APIsfal.aiImage, video, audio, and 3D APIs with per-output or per-second pricingPrepaid credits and per-model units need tracking
Production model inferenceFireworks AIServerless inference, batch jobs, dedicated GPUs, fine-tuning, and cached-token discountsNamed model rates, GPU utilization, batch timing, and cached-token behavior decide the real bill
Serverless Python/GPU appsModalPython jobs, web endpoints, queues, sandboxes, and per-second GPU billing without KubernetesRegion selection, non-preemptible execution, and steady 24/7 GPU load can change the economics
Cloud browser infrastructureBrowserbaseManaged Chromium sessions, web data APIs, Functions runtime, identity, Model Gateway, observability, Stagehand, and MCPBrowser sessions, Fetch/Extract calls, proxy bandwidth, model tokens, and agent loops need cost, timeout, and credential controls
Web data for agentsFirecrawlSearch, scrape, crawl, screenshots, Interact, and LLM-ready extraction in one APIEndpoint mix, crawl depth, screenshots, browser minutes, robots, terms, and retry behavior decide production fit
Agent tool/auth layerComposio1000+ app toolkits, user-scoped auth, session tools, and hosted MCP access for agent productsTool-call volume, OAuth scopes, write-action approvals, and multi-tenant identity separation need controls
Open-source LLM gatewayLiteLLMOne OpenAI-compatible interface across 100+ providers, with proxy routing, virtual keys, spend tracking, guardrails, MCP, and enterprise controlsGateway latency, log retention, fallback quality, enterprise-directory licensing, and model-provider bills need review
LLMOps plus gateway controlRespanTraces, metrics, evals, prompt management, monitors, spend limits, and an OpenAI-compatible gateway share one production traffic historyLive Team price, gateway latency, retention, data policy, and model-provider bills need confirmation
OpenTelemetry-first LLM observabilityOpenLITApache-2.0 self-hosting for traces, metrics, token and cost tracking, prompts, evals, dashboards, vector database monitoring, and GPU monitoringManaged cloud pricing is not public, and self-hosted telemetry operations stay with the team
OSS-to-hosted agent evalsOpikOpen-source or hosted agent traces, Test Suites, assertions, LLM-as-judge metrics, annotation, and production monitoringSpan volume, retention, additional-span pricing, and evaluator quality need budgets
Custom agent and safety evalsInspect AIPython and CLI eval tasks, datasets, scorers, tools, agents, Inspect View, and sandboxed runsModel calls, sandboxes, judge calibration, private data, and reviewer time remain buyer-owned
Runtime validation guardrailsGuardrails AIValidators, Hub installs, Pydantic outputs, on-fail policies, and input/output guards around LLM appsHosted pricing was not publicly verified; false positives and semantic correctness still need evals
Standard model benchmarksLM Evaluation Harness60+ academic benchmarks, leaderboard-style tasks, and HF, vLLM, API, SGLang, and local-server backendsBenchmarks do not replace private product evals; API and GPU costs can grow
OpenAI eval migrationOpenAI EvalsExisting OpenAI Evals users can inventory and migrate eval assets before read-only and shutdown datesExisting evals become read-only on October 31, 2026, and dashboard/API shutdown is scheduled for November 30, 2026
LLM and agent observabilityLangSmith, Arize Phoenix, Traceloop, LangWatch, or BraintrustLangSmith is LangChain-native operations; Phoenix and Traceloop are OpenTelemetry-aligned observability; LangWatch is open-source LLMOps; Braintrust is evals, datasets, scores, and release evidenceTrace volume, retention, eval/model spend, event meters, human review, usage meters, and acquisition/product-roadmap risk need limits
Enterprise AI reliabilityPatronus AIManaged evals, traces, datasets, guardrails, Percival-assisted eval creation, and Digital World Model researchConfirm whether the contract covers eval ops, Digital World Model access, services, retention, and security controls
Code-first LLM and RAG evalsDeepEval or RagasDeepEval is a broad LLM eval framework; Ragas is sharper for RAG metrics, synthetic test data, and cost-aware loopsTeams still own datasets, evaluator model costs, CI integration, and review workflow
Fine-tuning and specialized inferenceOpenPipeRequest logs, datasets, fine-tuning, DPO, evaluations, and hosted inference for cheaper specialized modelsNeeds clean training data, stable tasks, eval coverage, and rollback plans
Typed LLM function layerBAMLTyped function definitions, generated clients, robust parsing, tests, streaming inputs, and optional Boundary Studio tracesIt is a framework choice, not a finished SaaS purchase; model, hosting, CI, and observability costs remain separate
Structured output libraryInstructorMIT-licensed validated JSON, Pydantic-style schemas, retries, and provider adapters for LLM app codeValid schemas can still be wrong; retry cost, evals, and monitoring remain separate
Agent platform runtimeAgnoApache-2.0 SDK and AgentOS path for agents, teams, workflows, memory, knowledge, storage, traces, audit logs, and interfacesTool permissions, memory retention, model budgets, and control-plane value need proof
Conversational AI app shellChainlitApache-2.0 Python framework for chat UIs around prototypes, RAG, internal tools, and agent demosNot a managed support bot; auth, persistence, hosting, monitoring, and support need ownership
LLM gateway governancePortkey or HeliconeRouting, logs, prompts, guardrails, keys, caching, budgets, fallback, and traffic policyGateway latency, recorded logs, retention, provider spend, and guardrail ownership need testing
Agent memory layerZep or Mem0Managed or self-hosted persistent memories across users, sessions, agents, and temporal context graphsMemory extraction, deletion, consent, vector/database ownership, and privacy review decide production safety
AI security eval harnesspromptfooRed teaming, vulnerability scanning, guardrails, model security, MCP proxy, local evals, and enterprise/on-premise testingTests need real targets, remediation owners, false-positive review, and runtime controls
Search API for agentsTavilySearch, extract, crawl, map, and research APIs for agent and RAG systemsCredit burn depends on depth, extraction volume, crawls, research mode, retries, and loop behavior
RAG and agents over private dataLlamaIndex or HaystackLlamaIndex is strong for context augmentation, indexing, LlamaCloud, and LlamaParse; Haystack is strong for open-source components, pipelines, document stores, and agentsFramework cost, parsing credits, embeddings, vector storage, hosting, retrieval evals, and access controls are separate
Prompt and program optimizationDSPySignatures, modules, metrics, and optimizers help tune repeated LLM tasks from examplesWeak metrics, small datasets, and optimizer token spend can make the wrong behavior look better
Speech and voice infrastructureDeepgramSTT, TTS, audio intelligence, and voice-agent APIsVoice minutes, channels, model choice, and LLM orchestration affect cost
Model discovery and endpointsHugging FaceModel cards, datasets, Spaces, Inference EndpointsLicense and safety checks stay with the builder
Production retrievalPinecone, Weaviate, or QdrantManaged or open vector search for RAGIndex design and embedding cost matter as much as database pricing

How to Choose

  • Model routing: Pick OpenRouter when you need one OpenAI-compatible API across many providers.
  • Direct LLM APIs: Pick OpenAI, Claude, or Gemini when native features, procurement, and provider-specific controls matter.
  • Cost and latency: Pick Mistral AI, OpenRouter, or Groq when you can benchmark quality against real prompts and need tighter unit economics.
  • Open-model infrastructure: Pick Together AI when you need hosted inference, fine-tuning, dedicated endpoints, code sandboxes, and GPU capacity for open-model products. The June 25 check separates serverless model-token pricing from dedicated inference and GPU clusters, so benchmark your actual traffic before assuming a single “open model” unit cost.
  • Model catalog and experiments: Pick Hugging Face for discovery, datasets, model cards, demos, Spaces, ZeroGPU, and endpoints.
  • Media and community models: Pick Replicate when the job is running image, video, audio, or custom models by API. The June 9 check confirms buyers should model public output-priced examples separately from hardware-time runs and private deployments that can bill while idle.
  • Fast media APIs: Pick fal.ai when successful-output billing, image/video/audio/3D endpoints, and fast experimentation matter.
  • Production inference: Pick Fireworks AI when hosted model APIs, batch inference, dedicated GPU deployments, and fine-tuning are more important than a polished chatbot UI.
  • Browser automation: Pick Browserbase when an AI agent, scraper, QA runner, or workflow needs managed browsers, Search/Fetch, Functions runtime, identity, observability, Model Gateway, and Stagehand-style automation.
  • Web data extraction: Pick Firecrawl when the agent needs pages transformed into markdown or structured data more than it needs a long-running browser session.
  • Agent tool access: Pick Composio when the hard problem is app actions, user authentication, session tools, MCP access, and connector maintenance.
  • LLM gateways: Pick LiteLLM when the team wants a self-hosted OpenAI-compatible gateway, routing, virtual keys, budgets, guardrails, and provider fallback before model calls reach apps.
  • Managed LLMOps and gateway: Pick Respan when traces, evals, prompts, monitors, spend limits, and gateway routing should share the same production traffic history.
  • OpenTelemetry-first observability: Pick OpenLIT when the team wants Apache-2.0 self-hosted LLM telemetry and is ready to own collectors, storage, retention, and dashboards.
  • Agent eval operations: Pick Opik when agent traces, Test Suites, LLM-as-judge metrics, and hosted span retention are the buying reason.
  • Custom eval frameworks: Pick Inspect AI when coding, agent, safety, reasoning, knowledge, behavior, or multimodal evals need to live in Python tasks with tools, scorers, and sandboxes.
  • Runtime guardrails: Pick Guardrails AI when validated structured outputs, reusable validators, and on-fail policies need to run before downstream actions trust model output.
  • Model benchmarks: Pick LM Evaluation Harness when standard academic or leaderboard-style benchmark comparability is the job.
  • OpenAI eval migration: Treat OpenAI Evals as a migration item because hosted evals become read-only on October 31, 2026 and shut down on November 30, 2026.
  • Agent observability: Pick LangSmith when LangChain or LangGraph production work needs traces, evals, prompt workflows, deployment, and usage controls in one managed lane.
  • AI evals: Pick Braintrust when release quality depends on datasets, experiments, prompt comparisons, scored outputs, and review workflows.
  • OpenTelemetry AI observability: Pick Arize Phoenix when traces, prompt iteration, evals, datasets, and experiments should sit close to engineering instrumentation.
  • Enterprise AI reliability: Pick Patronus AI when managed evals, traces, datasets, guardrails, and frontier agent-simulation research are part of the buying question.
  • OpenTelemetry LLM observability: Pick Traceloop when OpenLLMetry instrumentation, quality checks, alerts, and ServiceNow procurement alignment matter.
  • Open-source LLMOps: Pick LangWatch when the team wants traces, evals, datasets, DSPy optimization, and self-hosting in one open-source product surface.
  • Code-first LLM evals: Pick DeepEval for broad LLM, RAG, agent, safety, tracing, and CI evals in code, and Ragas for RAG-specific metrics, synthetic test data, and experiments in Python.
  • Fine-tuning: Pick OpenPipe when logs can become datasets and specialized models can reduce cost or latency.
  • Typed LLM calls: Pick BAML when structured outputs, generated clients, robust parsing, and tests should be part of the app code contract.
  • Gateway governance: Pick Portkey when live AI traffic needs routing, keys, prompts, guardrails, caching, logs, budgets, and retention policy in one control layer.
  • Agent memory: Pick Zep or Mem0 when the product needs explicit, persistent memory rather than only retrieved documents.
  • Red-team testing: Pick promptfoo when the team needs local evals, jailbreak proxy review, and security evidence before launch.
  • Agent search API: Pick Tavily when the app needs web search, extraction, crawl, mapping, and research endpoints rather than a browser session.
  • RAG frameworks: Pick LlamaIndex when context augmentation and agents over private data are the job; pick Haystack when open-source components, pipelines, document stores, and reusable LLM app architecture are the job.
  • LLM program optimization: Pick DSPy when a repeated LLM task has examples and metrics that can guide optimization.
  • Agent platform runtime: Pick Agno when developers want an open-source AgentOS-style stack for agents, teams, workflows, memory, knowledge, traces, audit logs, and interfaces.
  • Structured outputs: Pick Instructor when the problem is validated JSON or typed LLM output in application code.
  • Conversational AI interface: Pick Chainlit when a Python RAG or agent workflow needs a quick chat UI for prototypes or internal tools.
  • Speech APIs: Pick Deepgram when speech-to-text, voice agents, or audio intelligence are infrastructure, not just creator utilities.
  • Serverless GPU apps: Pick Modal when you want Python jobs, endpoints, queues, sandboxes, and GPU workloads without Kubernetes. The June 25 check keeps Starter at $0 with $30/month credits, Team at $250/month plus compute with $100/month credits, B200 at $0.001736/sec, H200 at $0.001261/sec, H100 at $0.001097/sec, region multipliers at 1.5x to 1.75x, non-preemptible execution at 3x, and B200+ as a compatibility route that can run on B200 or B300 while billing as B200.
  • Open-weight model family: Pick Llama when infrastructure needs self-hostable or provider-hosted open weights card at $0.11/M input and $0.34/M output, and Together Maverick at $0.27/M input and $0.85/M output. Treat provider-specific availability, context, and pricing as live checks.
  • Local model runtime: Pick LM Studio when developers need a desktop GUI plus native v1 REST API, OpenAI-compatible and Anthropic-compatible endpoints, MCP support, SDKs, CLI server control, and LM Link for Llama, Qwen, Mistral, and other open weights. LM Studio has been free for ordinary home and work use since its July 2025 terms change, while Enterprise is the sales-led route for SSO, model/MCP gating, and private collaboration.
  • Managed vector search: Pick Pinecone, backups, imports, and reranking before treating the database price as the whole retrieval bill.
  • Open vector databases and agent memory: Pick Weaviate or Qdrant when self-hosting optionality and control matter. The June 10 Weaviate check keeps Free, Flex from $45/month, Plus from $280/month, Premium from $400/month, Weaviate Embeddings at $0.025-$0.065 per 1M tokens, Query Agent at a free 1,000-request/month trial path or $30/org/month with 4,000 included requests, and Engram generally available as a managed memory/context service for agents. The June 25 Qdrant check keeps the Free Cloud testing tier at 1GB RAM and 4GB disk; Standard as usage-based production cloud; Premium as the enterprise-support tier; Hybrid/Private Cloud as the control-first path; and v1.18.2 as the latest release checked, with security fixes included in the release notes.

Money Pages To Keep Current

  • Best pay-as-you-go AI tools and APIs was refreshed June 27, 2026 to separate true metered API usage from flat subscriptions and keep OpenAI, Claude, Gemini, OpenRouter, Mistral, Groq, Replicate, fal, Deepgram, ElevenLabs, and Fish Audio pricing risk in one buyer path.
  • Best open source AI tools was refreshed June 27, 2026 for Ollama, LM Studio, Open WebUI, Llama, Mistral, DeepSeek, FLUX, Stable Diffusion, Whisper, and Hugging Face because open-model buyers often compare local control against hosted pay-as-you-go APIs.
  • Best AI tools for developers is the June 6 verified developer guide for separating Cursor, GitHub Copilot AI Credits, Claude Code shared limits/API credits, Codex token credits, Replit Agent, and Aider BYOK API costs.
  • A new OpenRouter vs direct APIs comparison would capture buyers choosing between a model router and direct OpenAI/Anthropic/Google contracts.
  • A new Replicate vs fal.ai comparison would capture image/video/API buyers choosing between broad model catalog and fast media-generation infrastructure.

Watchouts

Infrastructure tools are powerful because they hide messy systems. That can also hide cost and governance risk. Before standardizing, test real workloads, pin model routes where quality matters, model retry costs, and document what data can pass through each provider.

Do not publish infrastructure pages with old flat monthly subscription framing. The buyer question is usually total workload cost: input tokens, output tokens, cached tokens, web/search tools, video seconds, generated images, GPU runtime, voice minutes, channels, retries, and failed generations.

Sources

Category graph

AI Infrastructure & Model APIs decision hub

Build a comparison
Share LinkedIn
Spotted an error or want to share your experience with AI Infrastructure & Model APIs?

Every tool page is re-verified on a recurring cycle, and corrections land faster when readers flag them directly. If you spot a stale fact, a missing capability, or have used AI Infrastructure & Model APIs and want to share what worked or didn't, the editorial desk reviews every message sent through this form.

Email editorial@aipedia.wiki