Best free or budget
Hugging Face See Hugging Face plansAI Infrastructure & Model APIs
Updated June 28, 2026: compare OpenRouter, OpenAI API, Claude API, Gemini API, Mistral, Groq, Replicate, Modal, Browserbase, Firecrawl, Composio, LiteLLM, Respan, OpenLIT, Opik, Inspect AI, Guardrails AI, LM Evaluation Harness, OpenAI Evals, LlamaIndex, Haystack, Agno, Instructor, Chainlit, LangSmith, Braintrust, Patronus AI, DeepEval, Traceloop, Arize Phoenix, Ragas, OpenPipe, LangWatch, BAML, DSPy, Portkey, Zep, promptfoo, Mem0, Tavily, Deepgram, vector databases, and governance tradeoffs.
Free tier (25+ models, 50 req/day) · Pay-as-you-go (5.5% platform fee on 400+ models) · Enterprise custom
Best model router
OpenRouter
Unified LLM API for hundreds of models, with OpenAI-compatible requests, provider routing, fallbacks, app attribution, and per-model token pricing.
Editorial · no paid placements
Quick paths
Best hosted model catalog
Replicate See Replicate plansBuyer path
Source-backed shortlist
- Source
- Registered source
- Freshness
- Current
- Confidence
- High confidence
Best local or open-model starter
lm-studioLM Studio is the cleanest first stop when the buyer wants local model testing, a desktop workflow, and an OpenAI-compatible local API before choosing hosted inference.
- Source
- Registered source
- Freshness
- Current
- Confidence
- High confidence
- Verified
Best web data API for agents
FirecrawlFirecrawl is the better shortlist when a product needs search, scrape, crawl, structured extraction, screenshots, and LLM-ready markdown rather than full browser-session ownership.
- Source
- Registered source
- Freshness
- Current
- Confidence
- High confidence
- Verified
Best tool and auth layer for agents
ComposioComposio fits AI products that need app toolkits, managed user authentication, session tools, and hosted MCP access without building every connector internally.
- Source
- Registered source
- Freshness
- Current
- Confidence
- High confidence
- Verified
Best open-source LLM gateway
LiteLLMLiteLLM is the better shortlist when teams need one OpenAI-compatible interface across 100+ providers, with routing, virtual keys, spend tracking, guardrails, MCP, and enterprise gateway controls.
- Source
- Registered source
- Freshness
- Current
- Confidence
- High confidence
- Verified
Best framework for agents over data
LlamaIndexLlamaIndex fits teams building RAG, context augmentation, workflows, and agents over private or domain-specific data, with a managed LlamaCloud/LlamaParse route when parsing and retrieval should be outsourced.
- Source
- Registered source
- Freshness
- Current
- Confidence
- High confidence
- Verified
Best open-source RAG pipeline framework
HaystackHaystack is the better shortlist when developers want Apache-2.0 components, pipelines, document stores, agents, tools, and retrieval systems with an optional deepset managed-platform route.
- Source
- Registered source
- Freshness
- Current
- Confidence
- High confidence
- Verified
Best LangChain-native observability layer
LangSmithLangSmith is the better shortlist when LangChain or LangGraph teams need hosted traces, evals, prompt workflows, monitoring, deployment, sandboxes, Fleet, and Engine controls in one platform.
- Source
- Registered source
- Freshness
- Current
- Confidence
- High confidence
- Verified
Best evals-first release control
BraintrustBraintrust is the better shortlist when teams need datasets, experiments, traces, scores, prompt testing, monitoring, and human review tied to release decisions.
- Source
- Registered source
- Freshness
- Current
- Confidence
- High confidence
- Verified
Best open-source AI observability platform
Arize PhoenixArize Phoenix is the better shortlist when OpenTelemetry traces, evals, prompt iteration, datasets, and experiments need to sit close to engineering workflows.
- Source
- Registered source
- Freshness
- Current
- Confidence
- High confidence
- Verified
Best OpenTelemetry-first LLM observability stack
OpenLITOpenLIT is the better shortlist when engineers want Apache-2.0 self-hosted traces, metrics, token and cost tracking, prompt workflows, evals, dashboards, GPU monitoring, and OpenTelemetry alignment.
- Source
- Registered source
- Freshness
- Current
- Confidence
- High confidence
- Verified
Best OSS-to-hosted agent eval platform
OpikOpik fits teams that want open-source or hosted agent traces, Test Suites, assertions, LLM-as-judge metrics, annotation, production monitoring, and Comet-hosted span retention.
- Source
- Registered source
- Freshness
- Current
- Confidence
- High confidence
- Verified
Best custom agent and safety eval framework
Inspect AIInspect AI fits teams that need code-defined evaluations for coding, agentic tasks, reasoning, knowledge, behavior, multimodal understanding, tools, scorers, and sandboxed runs.
- Source
- Registered source
- Freshness
- Current
- Confidence
- High confidence
- Verified
Best runtime validation guardrails
Guardrails AIGuardrails AI fits teams that need reusable validators, input and output guards, Pydantic-style structured data, on-fail policies, and Guardrails Hub installs around LLM apps.
- Source
- Registered source
- Freshness
- Current
- Confidence
- High confidence
- Verified
Best standardized benchmark harness
LM Evaluation HarnessLM Evaluation Harness fits model teams that need reproducible academic and leaderboard-style benchmark runs across local models, hosted APIs, vLLM, SGLang, Hugging Face, and other backends.
- Source
- Registered source
- Freshness
- Current
- Confidence
- High confidence
- Verified
Best code-first RAG evaluation framework
RagasRagas fits developer teams that want open-source metrics, synthetic test data, experiments, and cost-aware eval loops in code rather than a hosted dashboard first.
- Source
- Registered source
- Freshness
- Current
- Confidence
- High confidence
- Verified
Best fine-tuning workflow for cheaper specialized models
OpenPipeOpenPipe is the better shortlist when production request logs can become datasets, fine-tunes, DPO runs, evaluations, and hosted inference for cost or latency reduction.
- Source
- Registered source
- Freshness
- Current
- Confidence
- High confidence
- Verified
Best open-source LLMOps platform
LangWatchLangWatch fits teams that want traces, evaluations, datasets, AI gateway workflows, DSPy optimization, self-hosting, and monitoring in one open-source LLMOps surface.
- Source
- Registered source
- Freshness
- Current
- Confidence
- High confidence
- Verified
Best combined LLMOps and gateway challenger
RespanRespan is the better shortlist when a team wants traces, metrics, evals, prompt management, monitors, spend limits, and an OpenAI-compatible gateway tied to the same production traffic.
- Source
- Registered source
- Freshness
- Current
- Confidence
- High confidence
- Verified
Best enterprise AI reliability lab
Patronus AIPatronus AI fits teams that need managed evals, traces, datasets, guardrails, Percival-assisted eval creation, and current Digital World Model research around agent simulation.
- Source
- Registered source
- Freshness
- Current
- Confidence
- High confidence
- Verified
Best open-source LLM evaluation framework
DeepEvalDeepEval is the better shortlist when developer-owned LLM, RAG, agent, chatbot, image, safety, and CI evals should live in code before a hosted quality platform is added.
- Source
- Registered source
- Freshness
- Current
- Confidence
- High confidence
- Verified
Best OpenTelemetry LLM observability layer
TraceloopTraceloop fits teams that want OpenLLMetry, OpenTelemetry traces, quality checks, prompt changes, real-time alerts, and a ServiceNow AI Control Tower path around agent runtime evidence.
- Source
- Registered source
- Freshness
- Current
- Confidence
- High confidence
- Verified
Best LLM gateway governance layer
PortkeyPortkey fits production AI teams that need routing, observability, prompt management, guardrails, API key control, budgets, caching, and enterprise gateway policy.
- Source
- Registered source
- Freshness
- Current
- Confidence
- High confidence
- Verified
Best red-team and eval harness
promptfoopromptfoo is the sharper first stop when the infrastructure risk is jailbreak testing, vulnerability scanning, guardrails, model security, MCP exposure, or repeatable local evals.
- Source
- Registered source
- Freshness
- Current
- Confidence
- High confidence
- Verified
All tools in AI Infrastructure & Model APIs
- 1
Hugging Face Open AI collaboration hub for models, datasets, Spaces, inference endpoints, evaluations, and enterprise ML workflows. - 2 LiteLLM Open-source LLM gateway and Python SDK for one OpenAI-compatible interface across 100+ model providers, with routing, virtual keys, spend tracking, guardrails, MCP, and enterprise controls.
- 3 promptfoo Open-source LLM evaluation, red teaming, vulnerability scanning, guardrails, model security, MCP proxy, code scanning, and enterprise AI security testing.
- 4 LlamaIndex Open-source framework and managed LlamaCloud stack for building LLM agents over private data, RAG, document parsing, extraction, indexing, retrieval, workflows, and context augmentation.
- 5 Arize Phoenix Open-source AI observability, tracing, evaluation, prompt engineering, experiments, and Arize AX hosting for teams improving LLM systems.
- 6 Braintrust AI evaluation, tracing, prompt playground, datasets, experiments, monitoring, human review, and observability infrastructure for teams shipping LLM products.
- 7
Inspect AI MIT-licensed evaluation framework from the UK AI Security Institute and Meridian Labs for coding, agent, reasoning, knowledge, behavior, and multimodal model evals. - 8 LangSmith LangChain's hosted agent and LLM observability platform for tracing, monitoring, evaluation, prompt workflows, deployment, sandboxes, Fleet agents, and Engine optimization.
- 9
LM Evaluation Harness MIT-licensed EleutherAI framework for running standardized language-model benchmarks across local models, APIs, vLLM, SGLang, Hugging Face, and leaderboard tasks. - 10
Modal Serverless cloud for Python, GPUs, jobs, web endpoints, sandboxes, queues, and AI apps that should scale without managing infrastructure. - 11
Portkey LLM gateway, observability, prompt management, routing, guardrails, governance, caching, and cost controls for production AI applications. - 12
Together AI AI infrastructure platform for serverless inference, dedicated GPU deployments, fine-tuning, code sandboxes, and open-model training workflows. - 13
Weaviate Open-source vector database and managed cloud for RAG, semantic search, hybrid search, multi-tenancy, embeddings, and AI-native retrieval. - 14 DeepEval Open-source LLM evaluation framework from Confident AI for metrics, test cases, RAG evals, agent evals, tracing, datasets, and CI-friendly quality gates.
- 15 Haystack Apache-2.0 AI orchestration framework from deepset for production LLM apps, RAG systems, agents, multimodal search, reusable components, pipelines, tools, and document stores.
- 16
OpenRouter Unified LLM API for hundreds of models, with OpenAI-compatible requests, provider routing, fallbacks, app attribution, and per-model token pricing. - 17
Pinecone Managed vector database for semantic search, hybrid search, RAG, recommendations, Pinecone Assistant, and production AI retrieval workloads. - 18
Qdrant Open-source vector database written in Rust, with managed cloud, Free/Standard/Premium tiers, hybrid/private cloud options, metadata filtering, payload indexes, and RAG-ready retrieval. - 19 Ragas Open-source evaluation framework for LLM apps, RAG systems, metrics, synthetic test data, experiments, and cost-aware eval loops.
- 20
Replicate Developer platform for running open and hosted AI models by API, with official models, community models, custom deployments, and usage-based pricing. - 21
LangWatch Open-source LLMOps platform for traces, evaluations, datasets, AI gateway workflows, DSPy optimization, self-hosting, and monitoring. - 22 Mem0 Memory layer for AI agents that persists user, session, and agent context across conversations, with a managed Platform and Apache-2.0 open-source self-hosting path.
- 23
OpenLIT Apache-2.0, OpenTelemetry-native LLM observability platform for traces, metrics, costs, prompts, evals, dashboards, and GPU monitoring. - 24 OpenPipe Fine-tuning, request logging, datasets, evaluations, DPO, and hosted inference for turning expensive prompts into cheaper specialized models.
- 25 Opik Open-source and hosted AI observability and evaluation platform from Comet for agent traces, test suites, LLM-as-judge metrics, and production monitoring.
- 26 Patronus AI AI evaluation and simulation infrastructure for LLM apps, agent debugging, evaluators, traces, datasets, prompts, guardrails, and Digital World Models.
- 27
Traceloop OpenTelemetry-based LLM observability and evaluation platform built on OpenLLMetry for traces, quality checks, prompt management, experiments, and enterprise AI monitoring. - 28 Zep Production agent memory and context engineering platform with temporal context graphs, credit-based plans, hosted cloud, BYOK, BYOC, and Graphiti open-source context graph work.
- 29 Composio Tool-calling and MCP infrastructure for AI agents, with 1000+ app toolkits, managed authentication, session tools, hosted MCP URLs, and usage-based pricing.
- 30
Guardrails AI Apache-2.0 guardrails framework and Hub for validating, structuring, and quality-controlling LLM inputs and outputs with reusable validators. - 31
Firecrawl Web data API for AI agents that search, scrape, crawl, parse, monitor, and interact with pages, then return LLM-ready markdown, HTML, screenshots, or structured data. - 32
Respan LLM engineering platform, formerly Keywords AI, for observability, evals, prompt management, and an OpenAI-compatible gateway across many model providers. - 33
OpenAI Evals OpenAI evaluation framework and platform for testing prompts, models, graders, and LLM app regressions, with a scheduled platform shutdown on November 30, 2026. - 34 DSPy MIT-licensed framework from Stanford for programming, optimizing, and evaluating language-model systems with signatures, modules, metrics, optimizers, agents, and structured inputs/outputs.
- 35 BAML Apache-2.0 language and toolchain from BoundaryML for typed LLM functions, generated clients, structured outputs, robust parsing, tests, streaming, multimodal inputs, and Boundary Studio traces.
- 36
Browserbase Cloud browser infrastructure for agents, scraping, QA automation, and web data workflows that need managed Chromium, Fetch, Search, identity, runtime, and observability. - 37
Instructor MIT-licensed structured-output library for getting validated JSON from LLMs with Pydantic-style models, retries, and provider adapters. - 38 Pydantic AI MIT-licensed Python agent framework from the Pydantic team, built around typed agents, structured outputs, tools, dependencies, MCP, evals, graph workflows, and Logfire observability.
- 39 Agno Apache-2.0 agent platform SDK and AgentOS control plane for building, running, observing, and managing production agent systems in your own stack.
- 40 Mirascope MIT-licensed provider-agnostic SDKs for typed LLM calls, tools, structured outputs, streaming, agents, and OpenTelemetry-friendly observability.
- 41 Outlines Apache-2.0 structured-generation library from .txt for constraining LLM output to JSON Schema, regex, grammars, and typed schemas during generation.
- 42 Tavily Real-time search, extract, crawl, map, and research APIs for AI agents and RAG workflows, priced by API credits.
- 43 Dify Open-source platform for building AI apps, agents, chatbots, workflows, and RAG systems, with Dify Cloud plus self-hosted Community, Premium, and Enterprise routes.
- 44
Flowise Open-source visual builder for AI agents and LLM workflows, with chatflows, agentflows, assistants, RAG pipelines, evaluations, tracing, teams, and self-hosting. - 45 Chainlit Apache-2.0 Python framework for building conversational AI interfaces, prototypes, and internal chat apps around LLM workflows.
Quick Decision
AI infrastructure tools sit underneath the apps people see. They route model calls, host open models, run GPU workloads, store embeddings, power RAG, transcribe audio, generate media, and help teams compare cost, latency, quality, and control without rebuilding the stack every month.
This category is for developer and platform buyers. If the user is choosing a chatbot, start with AI Chatbots. If the team is shipping an AI product, agent, retrieval layer, or model-backed workflow, this is the better lane.
The late-May infrastructure update is agent control. CoreWeave’s training-to-inference loop pushes traces, evals, RL, inference, and W&B tooling into one reliability story. OpenAI’s Rosalind Biodefense trusted-access expansion shows that specialist frontier models may ship as gated capability programs. Sysdig’s LLM-agent intrusion report makes runtime telemetry and least-privilege design part of infrastructure buying, not only security cleanup.
The June 3 update widens that control story. Microsoft Build put Work IQ and Foundry around enterprise agents; GitHub made the Copilot SDK generally available while AI Credits became the agent-usage meter; NVIDIA pushed enterprise agents, Cosmos 3, open physical-AI agent skills, Alpamayo 2 Super, RTX Spark, and DGX Station for Windows; Postman launched AI Engineer for API work; RelationalAI moved agentic decision intelligence deeper into Snowflake; 7AI kept security agents in the proactive-hunting lane; and the White House AI cybersecurity order put advanced AI cyber capability into public-sector and critical-infrastructure policy. Infrastructure buyers should evaluate agent stacks by context access, runtime isolation, traces, evals, spend controls, simulation/data pipelines, local-vs-cloud compute, and write-action approvals.
The June 24 update keeps model availability as a first-class infrastructure risk. Claude Fable/Mythos access is still not a normal public self-serve buyer route, GPT-5.2 is retired from ChatGPT, and OpenAI faces reported state-AG scrutiny. Direct frontier API buyers should now document the exact model route in production, the fallback route if a model is suspended, account-specific, preview-only, or retired, the retention policy for that model class, the staff/client access exposure for restricted routes, and the legal/privacy review path for sensitive users. The AI Model Availability & Churn Tracker is now the canonical AiPedia surface for these app/API/router distinctions.
The June 16 infrastructure update is governed data agents. Google Cloud’s data-agent rollout puts Conversational Analytics, Data Engineering Agent, Looker agents, Gemini Enterprise data access, Data Agent Kit, Managed MCP Servers for Databases are GA, while many of the more ambitious analytics, Looker, Gemini Enterprise, and commerce routes remain preview. Infrastructure teams should evaluate these by IAM scope, roles/mcp.toolUser, service permissions, separate production identities, SQL verification, BigQuery spend limits, job labels, audit logging, Model Armor payload logging, and GA-versus-preview fallback plans.
Use OpenRouter when you need one API across many model providers. The current pricing page lists pay-as-you-go access to 400+ models and 60+ providers, with budget controls, activity logs, prompt caching, preferred vendor selections, and model-priced token billing. Its May 27 funding signal makes the category clearer: routing, fallback, governance, and spend visibility are becoming production infrastructure, not just developer convenience.
Use direct vendor APIs when native features matter. OpenAI API is the default direct route for broad multimodal app work. Claude API is the direct route for long reasoning, writing, code, and document workflows. Gemini API inputs, or Veo video generation are part of the product. The June 22 Gemini recheck keeps Gemini 3.5 Flash pricing mode-specific: standard, batch/flex, priority, grounding, tools, and media rows need separate cost modeling.
Use Mistral AI or Groq when price/performance, open-model strategy, European infrastructure, or low-latency inference matters. The June 24 Groq check adds Qwen 3.6 27B, Kimi buyer math. The June 24 Mistral check keeps the timeline and cost model honest: Mistral 3 officially launched on December 2, 2025, while Medium 3.5’s model-card date is April 28, 2026. Current Mistral pricing lists Large 3 at $0.50/M input and $1.50/M output, Medium 3.5 at $1.50/M and $7.50/M, and Small 4 at $0.10/M and $0.30/M, but the Small 4 model card still lists $0.15/M and $0.60/M, and the pricing FAQ still uses a generic Mistral Large $2/$6 example. Benchmark real prompts, confirm the live Studio quote, and pin exact model IDs before switching because model quality, output length, retries, aliases, and source drift change the bill.
Use Replicate or fal.ai when the job is hosted image, video, audio, 3D, or custom-model inference. The June 9 Replicate check keeps it strongest as a broad model catalog and custom-model deployment layer: public models may bill by hardware time or by input/output, while most private deployments bill setup, idle, and active time unless they are labeled fast-booting fine-tunes. fal is stronger when successful-output billing and fast media APIs are the buyer problem; the June 2 check keeps prepaid credits, queue behavior, failed-output billing, and the 50% batch discount as the key pricing details to model.
Use Fireworks AI when the workload is production inference over open or commercial models., cached-token discounts, batch jobs, dedicated GPU deployments, fine-tuning, and B200/B300 capacity are the actual purchase.
Use Browserbase when the infrastructure problem is web interaction.. It belongs here when agents need reliable browser sessions, Fetch/Extract, replay, and model routing rather than just another LLM.
Use Firecrawl when the infrastructure problem is web data for agents. The June 28 check adds it as the cleaned-content counterpart to Browserbase: search, scrape, crawl, structured extraction, screenshots, Interact, and LLM-ready markdown in one credit-based API. Model the exact endpoint mix before buying because crawls, screenshots, browser minutes, and retries can change the bill quickly.
Use Composio when agents need app actions and user-scoped auth. The June 28 check keeps it in the tool-calling infrastructure lane: 1000+ app toolkits, managed authentication, native session tools, hosted MCP URLs, and usage-based tool-call pricing. It is a developer tool layer, not a no-code workflow canvas.
Use LiteLLM when LLM traffic needs an open-source gateway. The June 28 check adds LiteLLM/proxy lane for one OpenAI-compatible interface across 100+ providers, with routing, virtual keys, spend tracking, budgets, guardrails, and Enterprise controls such as SSO, logs, fallback quality, and enterprise-directory licensing still need review.
Use LangSmith when agent observability, evals, and deployment control need a LangChain-native home. The June 28 check keeps Developer at $0/seat/month, Plus at $39/seat/month, and Enterprise custom, but the buyer risk is broader than seats: traces, extended retention, deployment runs, deployment uptime, Fleet, Engine, sandboxes, and outside model/API spend all need usage limits.
Use Braintrust when the infrastructure problem is eval discipline. Braintrust belongs here when datasets, experiments, prompt comparisons, traces, scores, monitoring, and human review need to decide whether a model, prompt, retrieval, or agent change is safe to ship. Starter is useful for prototypes; Pro at $249/month is the first serious shared-team tier, but usage meters still matter.
Use Arize Phoenix when the infrastructure problem is OpenTelemetry-native AI observability. Phoenix belongs here for teams that need traces, evals, prompt iteration, datasets, and experiments tied to LLM app behavior. Self-hosted Phoenix is the open-source path; AX Pro at $50/month is the hosted small-team route, but span volume and Elastic License 2.0 limits need review.
Use OpenLIT when OpenTelemetry alignment is the first requirement. The June 28 check adds OpenLIT traces, metrics, logs, token and cost tracking, prompt workflows, evals, dashboards, vector database monitoring, and GPU monitoring. The self-hosted product is listed at $0 forever, while OpenLIT Cloud is coming soon with no public price verified.
Use Opik when agent traces and evals need an OSS-to-hosted path. Opik belongs here when teams want trace debugging, Test Suites, assertions, LLM-as-judge metrics, production monitoring, annotation queues, and Comet-hosted retention. Start with OSS or Free Cloud, then model Pro Cloud at $19/month plus span and retention meters before rollout.
Use Inspect AI when evals need to be code-defined and sandboxed. The June 28 check adds Inspect AI as the MIT-licensed UK AI Security Institute and Meridian Labs framework for coding, agentic, reasoning, knowledge, behavior, and multimodal evals. It supports datasets, agents, tools, scorers, Inspect View, a VS Code extension, external agents, MCP or custom tools, and sandbox routes. It is free software, but model calls, compute, sandboxes, storage, and review time still need budgets.
Use Guardrails AI when LLM outputs need runtime validation before downstream systems trust them. Guardrails adds reusable validators, Guards, input/output checks, Pydantic-style structured data, on-fail policies, and Guardrails Hub installs around app code. The Apache-2.0 framework is the validated route; public hosted or remote-validator pricing was not verified in the June 28 check.
Use LM Evaluation Harness when benchmark comparability matters. EleutherAI’s MIT framework is the standard-benchmark lane for 60+ academic benchmarks, hundreds of task variants, custom prompts and metrics, and backend routes across Hugging Face, vLLM, API models, OpenAI-compatible local servers, SGLang, OpenVINO, NeMo, Megatron-LM, and more. Pair it with private product evals before making release decisions.
Use OpenAI Evals only as a migration bridge. OpenAI’s deprecations page says Evals platform deprecation was announced June 3, 2026, existing evals become read-only on October 31, 2026, and the dashboard/API are scheduled to shut down on November 30, 2026. Existing users should inventory, export, or migrate eval coverage; new teams should choose a maintained eval stack instead.
Use LangWatch when the team wants open-source LLMOps with traces and evals. LangWatch sits between local eval frameworks and managed observability products: traces, evaluations, datasets, AI gateway workflows, DSPy optimization, self-hosting, and event-sourcing operations. Paid plans start from EUR59/month, but event and retention meters decide real cost.
Use Respan when LLMOps and gateway routing should share one traffic history. The June 28 check adds Respan as the active Keywords AI successor surface: traces, metrics, evals, prompt management, monitors, spend limits, and an OpenAI-compatible gateway sit in one product lane. Free is the first instrumentation route, Team is the shared paid route, and Enterprise is custom. Confirm the live Team base price, retention, data-handling terms, gateway latency, and provider bills before standardizing.
Use Patronus AI when AI reliability needs managed evals plus frontier simulation research. Patronus covers evaluators, experiments, datasets, comparisons, traces, prompts, annotations, guardrails, and Percival-assisted eval creation, while its current company positioning also highlights Digital World Models for simulating agent actions in digital workflows. Confirm whether procurement is buying eval operations, Digital World Model access, or services.
Use DeepEval when LLM evaluation should live in code. DeepEval evals, agent evals, chatbot tests, safety checks, tracing, and CI gates. Confident AI is the hosted platform route when teams need shared evaluation, observability, red teaming, governance, and security controls.
Use Traceloop when OpenTelemetry alignment is the observability requirement. Traceloop and OpenLLMetry instrument LLM apps with OpenTelemetry traces, quality checks, alerts, prompt and model change testing, and IDE reruns. The ServiceNow acquisition path can help enterprise buyers, but roadmap and procurement details need live confirmation.
Use Ragas when RAG and LLM evaluation should live in code. Ragas is the open-source framework lane for metrics, synthetic test data, experiments, and cost-aware evaluation loops. It is free as a framework, but evaluator model calls, embeddings, generated test data, and human review still need a budget.
Use OpenPipe when fine-tuning can lower cost or latency. OpenPipe belongs here when production logs can become datasets, fine-tuned models, DPO runs, evaluations, and hosted inference. Start only after the team has enough clean examples and baseline evals to prove a specialized model beats prompt-only changes.
Use Portkey when LLM traffic needs a gateway control plane. Portkey, governance, caching, and analytics. The Developer plan is explicitly not the production path; Production and Enterprise should be modeled around recorded logs, retention, overages, private hosting, compliance, and provider spend.
Use Zep when the infrastructure problem is production agent memory. Zep is the context-engineering lane for temporal context graphs, persistent user/session memory, and managed deployment options. It belongs beside Mem0 and vector databases, but buyers should decide whether they need memory extraction and graph context or only document retrieval.
Use promptfoo when AI security testing is part of the infrastructure gate. promptfoo is now positioned around open-source evals, red teaming, guardrails, model security, MCP proxy checks, code scanning, and enterprise/on-premise testing workflows. OpenAI announced a March 2026 acquisition agreement, and promptfoo says it is now part of OpenAI, so procurement should verify contract and data-handling paths.
Use Mem0 when the infrastructure problem is agent memory. Mem0 belongs beside vector databases and agent runtimes when an app needs persistent, scoped memories across users, sessions, and agents. The managed Platform is the faster route; the Apache-2.0 open-source route shifts vector database, model, embedder, hosting, deletion, and privacy controls back to the team.
Use Tavily when agents need real-time web search and extraction as an API. Tavily is closer to a search, extract, crawl, map, and research primitive than a browser infrastructure layer. Model basic versus advanced search, extract batches, mapping, crawl size, research mode, retries, and agent-loop behavior before production.
Use LlamaIndex when agents need structured context over private data. The June 28 check adds it as the framework lane for RAG, context augmentation, data connectors, indexes, retrieval, workflows, and agents over domain-specific data. The framework is MIT-licensed, but LlamaParse/LlamaCloud credits, model calls, embeddings, vector storage, ingestion jobs, and evals still need separate budgets.
Use Haystack when RAG and agent apps should be built as open-source pipelines. Haystack belongs here for Apache-2.0 components, pipelines, document stores, tools, agents, multimodal search, and reusable LLM app architecture. deepset AI Platform is the managed route when deployment, governance, testing, support, security, or dedicated resources matter.
Use DSPy when prompt behavior needs measurable optimization. DSPy is a coding-first framework, but it belongs in infrastructure shortlists when an LLM system needs signatures, modules, metrics, optimizers, and examples instead of hand-tuned prompt sprawl. Budget for model calls and data work before assuming optimization will lower costs.
Use Agno when agents need an owned AgentOS-style runtime. Agno is primarily an automation and agent-platform framework, but it belongs in infrastructure shortlists when teams need agents, teams, workflows, memory, knowledge, storage, traces, audit logs, interfaces, and control-plane management in their own stack. Free open source is the first route; Pro is $150/month for a live AgentOS control plane; Enterprise is custom.
Use Instructor when structured LLM outputs are the infrastructure boundary. Instructor is the MIT-licensed library lane for validated JSON, Pydantic-style schemas, retries, and provider adapters. It is useful when extraction, classification, enrichment, or tool arguments need typed outputs, but it still needs evals, retry budgets, and monitoring.
Use Chainlit when a Python LLM workflow needs a quick conversational interface. Chainlit is an Apache-2.0 framework for prototypes, internal chat tools, and demos around RAG or agent workflows. Treat it as a UI framework, not a managed chatbot platform, and verify maintainer/support route, auth, persistence, hosting, and observability before production use.
Use Deepgram when speech is infrastructure. Deepgram is a better fit for product teams adding STT, TTS, audio intelligence, or voice agents than for creators who only need a one-off transcript.
Use Hugging Face when model discovery, model cards, datasets, Spaces, and managed endpoints need to live in one open-AI collaboration surface. The June 23 pricing check keeps Pro at $9/month, Team at $20/user/month, Enterprise from $50/user/month, storage at $12/TB public and $18/TB private before volume discounts, ZeroGPU on RTX Pro 6000 Blackwell, and Inference Endpoints starting around low hourly CPU pricing.
Buyer Paths
| Buyer job | Start with | Why | Watch out |
|---|---|---|---|
| Multi-model LLM routing | OpenRouter | One API, many providers, spend controls, logs, routing | Router fees and provider policy choices still need governance |
| Direct frontier LLM API | OpenAI, Claude, or Gemini | Best when native model features, support, and procurement matter | Model access, retirements, legal/data governance, long context, outputs, tools, and video can change cost and risk quickly |
| Budget/open-model API | Mistral AI or Groq | Useful for cost-sensitive, latency-sensitive, and sovereignty-sensitive workloads | Requires benchmarking against your actual prompts, exact model IDs, and current model-card/pricing-page drift |
| Hosted model catalog | Replicate | Public, proprietary, and custom models without owning GPUs | Hardware-time, output-priced media, and private-model idle billing need separate cost modeling |
| Fast media APIs | fal.ai | Image, video, audio, and 3D APIs with per-output or per-second pricing | Prepaid credits and per-model units need tracking |
| Production model inference | Fireworks AI | Serverless inference, batch jobs, dedicated GPUs, fine-tuning, and cached-token discounts | Named model rates, GPU utilization, batch timing, and cached-token behavior decide the real bill |
| Serverless Python/GPU apps | Modal | Python jobs, web endpoints, queues, sandboxes, and per-second GPU billing without Kubernetes | Region selection, non-preemptible execution, and steady 24/7 GPU load can change the economics |
| Cloud browser infrastructure | Browserbase | Managed Chromium sessions, web data APIs, Functions runtime, identity, Model Gateway, observability, Stagehand, and MCP | Browser sessions, Fetch/Extract calls, proxy bandwidth, model tokens, and agent loops need cost, timeout, and credential controls |
| Web data for agents | Firecrawl | Search, scrape, crawl, screenshots, Interact, and LLM-ready extraction in one API | Endpoint mix, crawl depth, screenshots, browser minutes, robots, terms, and retry behavior decide production fit |
| Agent tool/auth layer | Composio | 1000+ app toolkits, user-scoped auth, session tools, and hosted MCP access for agent products | Tool-call volume, OAuth scopes, write-action approvals, and multi-tenant identity separation need controls |
| Open-source LLM gateway | LiteLLM | One OpenAI-compatible interface across 100+ providers, with proxy routing, virtual keys, spend tracking, guardrails, MCP, and enterprise controls | Gateway latency, log retention, fallback quality, enterprise-directory licensing, and model-provider bills need review |
| LLMOps plus gateway control | Respan | Traces, metrics, evals, prompt management, monitors, spend limits, and an OpenAI-compatible gateway share one production traffic history | Live Team price, gateway latency, retention, data policy, and model-provider bills need confirmation |
| OpenTelemetry-first LLM observability | OpenLIT | Apache-2.0 self-hosting for traces, metrics, token and cost tracking, prompts, evals, dashboards, vector database monitoring, and GPU monitoring | Managed cloud pricing is not public, and self-hosted telemetry operations stay with the team |
| OSS-to-hosted agent evals | Opik | Open-source or hosted agent traces, Test Suites, assertions, LLM-as-judge metrics, annotation, and production monitoring | Span volume, retention, additional-span pricing, and evaluator quality need budgets |
| Custom agent and safety evals | Inspect AI | Python and CLI eval tasks, datasets, scorers, tools, agents, Inspect View, and sandboxed runs | Model calls, sandboxes, judge calibration, private data, and reviewer time remain buyer-owned |
| Runtime validation guardrails | Guardrails AI | Validators, Hub installs, Pydantic outputs, on-fail policies, and input/output guards around LLM apps | Hosted pricing was not publicly verified; false positives and semantic correctness still need evals |
| Standard model benchmarks | LM Evaluation Harness | 60+ academic benchmarks, leaderboard-style tasks, and HF, vLLM, API, SGLang, and local-server backends | Benchmarks do not replace private product evals; API and GPU costs can grow |
| OpenAI eval migration | OpenAI Evals | Existing OpenAI Evals users can inventory and migrate eval assets before read-only and shutdown dates | Existing evals become read-only on October 31, 2026, and dashboard/API shutdown is scheduled for November 30, 2026 |
| LLM and agent observability | LangSmith, Arize Phoenix, Traceloop, LangWatch, or Braintrust | LangSmith is LangChain-native operations; Phoenix and Traceloop are OpenTelemetry-aligned observability; LangWatch is open-source LLMOps; Braintrust is evals, datasets, scores, and release evidence | Trace volume, retention, eval/model spend, event meters, human review, usage meters, and acquisition/product-roadmap risk need limits |
| Enterprise AI reliability | Patronus AI | Managed evals, traces, datasets, guardrails, Percival-assisted eval creation, and Digital World Model research | Confirm whether the contract covers eval ops, Digital World Model access, services, retention, and security controls |
| Code-first LLM and RAG evals | DeepEval or Ragas | DeepEval is a broad LLM eval framework; Ragas is sharper for RAG metrics, synthetic test data, and cost-aware loops | Teams still own datasets, evaluator model costs, CI integration, and review workflow |
| Fine-tuning and specialized inference | OpenPipe | Request logs, datasets, fine-tuning, DPO, evaluations, and hosted inference for cheaper specialized models | Needs clean training data, stable tasks, eval coverage, and rollback plans |
| Typed LLM function layer | BAML | Typed function definitions, generated clients, robust parsing, tests, streaming inputs, and optional Boundary Studio traces | It is a framework choice, not a finished SaaS purchase; model, hosting, CI, and observability costs remain separate |
| Structured output library | Instructor | MIT-licensed validated JSON, Pydantic-style schemas, retries, and provider adapters for LLM app code | Valid schemas can still be wrong; retry cost, evals, and monitoring remain separate |
| Agent platform runtime | Agno | Apache-2.0 SDK and AgentOS path for agents, teams, workflows, memory, knowledge, storage, traces, audit logs, and interfaces | Tool permissions, memory retention, model budgets, and control-plane value need proof |
| Conversational AI app shell | Chainlit | Apache-2.0 Python framework for chat UIs around prototypes, RAG, internal tools, and agent demos | Not a managed support bot; auth, persistence, hosting, monitoring, and support need ownership |
| LLM gateway governance | Portkey or Helicone | Routing, logs, prompts, guardrails, keys, caching, budgets, fallback, and traffic policy | Gateway latency, recorded logs, retention, provider spend, and guardrail ownership need testing |
| Agent memory layer | Zep or Mem0 | Managed or self-hosted persistent memories across users, sessions, agents, and temporal context graphs | Memory extraction, deletion, consent, vector/database ownership, and privacy review decide production safety |
| AI security eval harness | promptfoo | Red teaming, vulnerability scanning, guardrails, model security, MCP proxy, local evals, and enterprise/on-premise testing | Tests need real targets, remediation owners, false-positive review, and runtime controls |
| Search API for agents | Tavily | Search, extract, crawl, map, and research APIs for agent and RAG systems | Credit burn depends on depth, extraction volume, crawls, research mode, retries, and loop behavior |
| RAG and agents over private data | LlamaIndex or Haystack | LlamaIndex is strong for context augmentation, indexing, LlamaCloud, and LlamaParse; Haystack is strong for open-source components, pipelines, document stores, and agents | Framework cost, parsing credits, embeddings, vector storage, hosting, retrieval evals, and access controls are separate |
| Prompt and program optimization | DSPy | Signatures, modules, metrics, and optimizers help tune repeated LLM tasks from examples | Weak metrics, small datasets, and optimizer token spend can make the wrong behavior look better |
| Speech and voice infrastructure | Deepgram | STT, TTS, audio intelligence, and voice-agent APIs | Voice minutes, channels, model choice, and LLM orchestration affect cost |
| Model discovery and endpoints | Hugging Face | Model cards, datasets, Spaces, Inference Endpoints | License and safety checks stay with the builder |
| Production retrieval | Pinecone, Weaviate, or Qdrant | Managed or open vector search for RAG | Index design and embedding cost matter as much as database pricing |
How to Choose
- Model routing: Pick OpenRouter when you need one OpenAI-compatible API across many providers.
- Direct LLM APIs: Pick OpenAI, Claude, or Gemini when native features, procurement, and provider-specific controls matter.
- Cost and latency: Pick Mistral AI, OpenRouter, or Groq when you can benchmark quality against real prompts and need tighter unit economics.
- Open-model infrastructure: Pick Together AI when you need hosted inference, fine-tuning, dedicated endpoints, code sandboxes, and GPU capacity for open-model products. The June 25 check separates serverless model-token pricing from dedicated inference and GPU clusters, so benchmark your actual traffic before assuming a single “open model” unit cost.
- Model catalog and experiments: Pick Hugging Face for discovery, datasets, model cards, demos, Spaces, ZeroGPU, and endpoints.
- Media and community models: Pick Replicate when the job is running image, video, audio, or custom models by API. The June 9 check confirms buyers should model public output-priced examples separately from hardware-time runs and private deployments that can bill while idle.
- Fast media APIs: Pick fal.ai when successful-output billing, image/video/audio/3D endpoints, and fast experimentation matter.
- Production inference: Pick Fireworks AI when hosted model APIs, batch inference, dedicated GPU deployments, and fine-tuning are more important than a polished chatbot UI.
- Browser automation: Pick Browserbase when an AI agent, scraper, QA runner, or workflow needs managed browsers, Search/Fetch, Functions runtime, identity, observability, Model Gateway, and Stagehand-style automation.
- Web data extraction: Pick Firecrawl when the agent needs pages transformed into markdown or structured data more than it needs a long-running browser session.
- Agent tool access: Pick Composio when the hard problem is app actions, user authentication, session tools, MCP access, and connector maintenance.
- LLM gateways: Pick LiteLLM when the team wants a self-hosted OpenAI-compatible gateway, routing, virtual keys, budgets, guardrails, and provider fallback before model calls reach apps.
- Managed LLMOps and gateway: Pick Respan when traces, evals, prompts, monitors, spend limits, and gateway routing should share the same production traffic history.
- OpenTelemetry-first observability: Pick OpenLIT when the team wants Apache-2.0 self-hosted LLM telemetry and is ready to own collectors, storage, retention, and dashboards.
- Agent eval operations: Pick Opik when agent traces, Test Suites, LLM-as-judge metrics, and hosted span retention are the buying reason.
- Custom eval frameworks: Pick Inspect AI when coding, agent, safety, reasoning, knowledge, behavior, or multimodal evals need to live in Python tasks with tools, scorers, and sandboxes.
- Runtime guardrails: Pick Guardrails AI when validated structured outputs, reusable validators, and on-fail policies need to run before downstream actions trust model output.
- Model benchmarks: Pick LM Evaluation Harness when standard academic or leaderboard-style benchmark comparability is the job.
- OpenAI eval migration: Treat OpenAI Evals as a migration item because hosted evals become read-only on October 31, 2026 and shut down on November 30, 2026.
- Agent observability: Pick LangSmith when LangChain or LangGraph production work needs traces, evals, prompt workflows, deployment, and usage controls in one managed lane.
- AI evals: Pick Braintrust when release quality depends on datasets, experiments, prompt comparisons, scored outputs, and review workflows.
- OpenTelemetry AI observability: Pick Arize Phoenix when traces, prompt iteration, evals, datasets, and experiments should sit close to engineering instrumentation.
- Enterprise AI reliability: Pick Patronus AI when managed evals, traces, datasets, guardrails, and frontier agent-simulation research are part of the buying question.
- OpenTelemetry LLM observability: Pick Traceloop when OpenLLMetry instrumentation, quality checks, alerts, and ServiceNow procurement alignment matter.
- Open-source LLMOps: Pick LangWatch when the team wants traces, evals, datasets, DSPy optimization, and self-hosting in one open-source product surface.
- Code-first LLM evals: Pick DeepEval for broad LLM, RAG, agent, safety, tracing, and CI evals in code, and Ragas for RAG-specific metrics, synthetic test data, and experiments in Python.
- Fine-tuning: Pick OpenPipe when logs can become datasets and specialized models can reduce cost or latency.
- Typed LLM calls: Pick BAML when structured outputs, generated clients, robust parsing, and tests should be part of the app code contract.
- Gateway governance: Pick Portkey when live AI traffic needs routing, keys, prompts, guardrails, caching, logs, budgets, and retention policy in one control layer.
- Agent memory: Pick Zep or Mem0 when the product needs explicit, persistent memory rather than only retrieved documents.
- Red-team testing: Pick promptfoo when the team needs local evals, jailbreak proxy review, and security evidence before launch.
- Agent search API: Pick Tavily when the app needs web search, extraction, crawl, mapping, and research endpoints rather than a browser session.
- RAG frameworks: Pick LlamaIndex when context augmentation and agents over private data are the job; pick Haystack when open-source components, pipelines, document stores, and reusable LLM app architecture are the job.
- LLM program optimization: Pick DSPy when a repeated LLM task has examples and metrics that can guide optimization.
- Agent platform runtime: Pick Agno when developers want an open-source AgentOS-style stack for agents, teams, workflows, memory, knowledge, traces, audit logs, and interfaces.
- Structured outputs: Pick Instructor when the problem is validated JSON or typed LLM output in application code.
- Conversational AI interface: Pick Chainlit when a Python RAG or agent workflow needs a quick chat UI for prototypes or internal tools.
- Speech APIs: Pick Deepgram when speech-to-text, voice agents, or audio intelligence are infrastructure, not just creator utilities.
- Serverless GPU apps: Pick Modal when you want Python jobs, endpoints, queues, sandboxes, and GPU workloads without Kubernetes. The June 25 check keeps Starter at $0 with $30/month credits, Team at $250/month plus compute with $100/month credits, B200 at $0.001736/sec, H200 at $0.001261/sec, H100 at $0.001097/sec, region multipliers at 1.5x to 1.75x, non-preemptible execution at 3x, and B200+ as a compatibility route that can run on B200 or B300 while billing as B200.
- Open-weight model family: Pick Llama when infrastructure needs self-hostable or provider-hosted open weights card at $0.11/M input and $0.34/M output, and Together Maverick at $0.27/M input and $0.85/M output. Treat provider-specific availability, context, and pricing as live checks.
- Local model runtime: Pick LM Studio when developers need a desktop GUI plus native v1 REST API, OpenAI-compatible and Anthropic-compatible endpoints, MCP support, SDKs, CLI server control, and LM Link for Llama, Qwen, Mistral, and other open weights. LM Studio has been free for ordinary home and work use since its July 2025 terms change, while Enterprise is the sales-led route for SSO, model/MCP gating, and private collaboration.
- Managed vector search: Pick Pinecone, backups, imports, and reranking before treating the database price as the whole retrieval bill.
- Open vector databases and agent memory: Pick Weaviate or Qdrant when self-hosting optionality and control matter. The June 10 Weaviate check keeps Free, Flex from $45/month, Plus from $280/month, Premium from $400/month, Weaviate Embeddings at $0.025-$0.065 per 1M tokens, Query Agent at a free 1,000-request/month trial path or $30/org/month with 4,000 included requests, and Engram generally available as a managed memory/context service for agents. The June 25 Qdrant check keeps the Free Cloud testing tier at 1GB RAM and 4GB disk; Standard as usage-based production cloud; Premium as the enterprise-support tier; Hybrid/Private Cloud as the control-first path; and v1.18.2 as the latest release checked, with security fixes included in the release notes.
Money Pages To Keep Current
- Best pay-as-you-go AI tools and APIs was refreshed June 27, 2026 to separate true metered API usage from flat subscriptions and keep OpenAI, Claude, Gemini, OpenRouter, Mistral, Groq, Replicate, fal, Deepgram, ElevenLabs, and Fish Audio pricing risk in one buyer path.
- Best open source AI tools was refreshed June 27, 2026 for Ollama, LM Studio, Open WebUI, Llama, Mistral, DeepSeek, FLUX, Stable Diffusion, Whisper, and Hugging Face because open-model buyers often compare local control against hosted pay-as-you-go APIs.
- Best AI tools for developers is the June 6 verified developer guide for separating Cursor, GitHub Copilot AI Credits, Claude Code shared limits/API credits, Codex token credits, Replit Agent, and Aider BYOK API costs.
- A new
OpenRouter vs direct APIscomparison would capture buyers choosing between a model router and direct OpenAI/Anthropic/Google contracts. - A new
Replicate vs fal.aicomparison would capture image/video/API buyers choosing between broad model catalog and fast media-generation infrastructure.
Watchouts
Infrastructure tools are powerful because they hide messy systems. That can also hide cost and governance risk. Before standardizing, test real workloads, pin model routes where quality matters, model retry costs, and document what data can pass through each provider.
Do not publish infrastructure pages with old flat monthly subscription framing. The buyer question is usually total workload cost: input tokens, output tokens, cached tokens, web/search tools, video seconds, generated images, GPU runtime, voice minutes, channels, retries, and failed generations.
Sources
- OpenRouter pricing (verified 2026-05-27)
- OpenRouter Series B announcement (verified 2026-05-27)
- OpenAI API pricing (verified 2026-06-12)
- Claude API pricing (verified 2026-06-12)
- AiPedia late June 13 AI news update (verified 2026-06-13)
- AiPedia June 14 AI news desk (verified 2026-06-14)
- Anthropic Fable/Mythos access statement (verified 2026-06-14)
- OpenAI ChatGPT release notes (verified 2026-06-13)
- Gemini API pricing (verified 2026-06-22)
- Mistral AI pricing (verified 2026-06-15)
- Mistral Vibe product page (verified 2026-06-15)
- Mistral Vibe agent announcement (verified 2026-06-15)
- Mistral AI Now Summit 2026 (verified 2026-06-15)
- Mistral model docs (verified 2026-06-15)
- Mistral 3 launch post (verified 2026-06-15)
- Groq pricing (verified 2026-06-23)
- Replicate pricing (verified 2026-06-12)
- fal Model API pricing docs (verified 2026-06-12)
- Fireworks AI pricing (verified 2026-06-12)
- Fireworks billing FAQ (verified 2026-06-12)
- Fireworks inference documentation (verified 2026-06-12)
- Browserbase pricing (verified 2026-06-18)
- Browserbase changelog (verified 2026-06-18)
- Browserbase Browser explainer (verified 2026-06-18)
- Browserbase Model Gateway docs (verified 2026-06-18)
- Firecrawl pricing (verified 2026-06-28)
- Firecrawl introduction docs (verified 2026-06-28)
- Firecrawl scrape docs (verified 2026-06-28)
- Composio pricing (verified 2026-06-28)
- Composio docs (verified 2026-06-28)
- Composio authentication docs (verified 2026-06-28)
- LiteLLM docs (verified 2026-06-28)
- LiteLLM Enterprise docs (verified 2026-06-28)
- LiteLLM license (verified 2026-06-28)
- LlamaIndex framework docs (verified 2026-06-28)
- LlamaIndex pricing (verified 2026-06-28)
- LlamaIndex license (verified 2026-06-28)
- Haystack introduction docs (verified 2026-06-28)
- deepset AI Platform pricing (verified 2026-06-28)
- Haystack license (verified 2026-06-28)
- DSPy official site (verified 2026-06-28)
- DSPy program, don’t prompt guide (verified 2026-06-28)
- DSPy license (verified 2026-06-28)
- Respan pricing (verified 2026-06-28)
- Respan gateway docs (verified 2026-06-28)
- Respan eval docs (verified 2026-06-28)
- Agno pricing (verified 2026-06-28)
- Agno docs (verified 2026-06-28)
- Instructor docs (verified 2026-06-28)
- Instructor license (verified 2026-06-28)
- Chainlit docs (verified 2026-06-28)
- Chainlit GitHub repository (verified 2026-06-28)
- OpenLIT pricing (verified 2026-06-28)
- OpenLIT overview docs (verified 2026-06-28)
- OpenLIT license (verified 2026-06-28)
- Opik pricing (verified 2026-06-28)
- Opik tracing docs (verified 2026-06-28)
- Opik evaluation docs (verified 2026-06-28)
- Inspect AI docs (verified 2026-06-28)
- Inspect AI evals list (verified 2026-06-28)
- Inspect AI license (verified 2026-06-28)
- Guardrails quickstart (verified 2026-06-28)
- Guardrails validators docs (verified 2026-06-28)
- Guardrails license (verified 2026-06-28)
- LM Evaluation Harness README (verified 2026-06-28)
- LM Evaluation Harness license (verified 2026-06-28)
- OpenAI evals guide (verified 2026-06-28)
- OpenAI API deprecations (verified 2026-06-28)
- OpenAI Evals license (verified 2026-06-28)
- LangSmith observability (verified 2026-06-28)
- LangChain pricing (verified 2026-06-28)
- LangSmith usage and billing (verified 2026-06-28)
- Braintrust pricing (verified 2026-06-28)
- Braintrust docs index (verified 2026-06-28)
- Arize Phoenix docs index (verified 2026-06-28)
- Arize pricing (verified 2026-06-28)
- Ragas docs (verified 2026-06-28)
- Ragas cost analysis docs (verified 2026-06-28)
- OpenPipe docs index (verified 2026-06-28)
- OpenPipe pricing (verified 2026-06-28)
- LangWatch docs index (verified 2026-06-28)
- LangWatch pricing (verified 2026-06-28)
- Patronus AI docs (verified 2026-06-28)
- Patronus AI pricing (verified 2026-06-28)
- Patronus Digital World Model announcement (verified 2026-06-28)
- DeepEval metrics docs (verified 2026-06-28)
- Confident AI pricing (verified 2026-06-28)
- Traceloop docs (verified 2026-06-28)
- Traceloop pricing (verified 2026-06-28)
- Traceloop joins ServiceNow (verified 2026-06-28)
- BAML docs index (verified 2026-06-28)
- Portkey pricing (verified 2026-06-28)
- Portkey docs (verified 2026-06-28)
- Zep pricing (verified 2026-06-28)
- Zep docs concepts (verified 2026-06-28)
- promptfoo pricing (verified 2026-06-28)
- promptfoo docs (verified 2026-06-28)
- OpenAI promptfoo acquisition announcement (verified 2026-06-28)
- Mem0 Platform overview (verified 2026-06-28)
- Mem0 Platform vs Open Source (verified 2026-06-28)
- Mem0 pricing (verified 2026-06-28)
- Tavily API credits (verified 2026-06-28)
- Deepgram pricing (verified 2026-06-12)
- Together AI pricing (verified 2026-06-25)
- Hugging Face pricing (verified 2026-06-23)
- Hugging Face ZeroGPU docs (verified 2026-06-23)
- Modal pricing (verified 2026-06-25)
- Modal GPU docs (verified 2026-06-25)
- LM Studio (verified 2026-06-23)
- LM Studio developer docs (verified 2026-06-23)
- LM Studio desktop app terms (verified 2026-06-23)
- Llama official site (verified 2026-06-23)
- Together AI Llama pricing (verified 2026-06-23)
- Groq Llama 4 Scout model card (verified 2026-06-23)
- Groq Llama 4 Maverick model card (verified 2026-06-23)
- Pinecone pricing (verified 2026-06-25)
- Pinecone cost docs (verified 2026-06-25)
- Pinecone Assistant pricing and limits (verified 2026-06-25)
- Weaviate pricing (verified 2026-06-12)
- Weaviate Engram GA announcement (verified 2026-06-12)
- Qdrant pricing (verified 2026-06-25)
- Qdrant Cloud billing (verified 2026-06-25)
- Qdrant v1.18.2 release notes (verified 2026-06-25)
- CoreWeave autonomous agent improvement launch (verified 2026-05-31)
- OpenAI Rosalind Biodefense (verified 2026-05-31)
- Geordie AI Series A (verified 2026-05-31)
- Sysdig LLM-agent intrusion analysis (verified 2026-05-31)
- Microsoft Build 2026 Work IQ and Foundry agent stack (verified 2026-06-12)
- Google Cloud data agents announcement (verified 2026-06-16)
- Google Cloud Data Engineering Agent docs (verified 2026-06-16)
- Google Cloud MCP servers docs (verified 2026-06-16)
- Google Cloud Conversational Analytics docs (verified 2026-06-16)
- GitHub Copilot SDK GA (verified 2026-06-12)
- NVIDIA enterprise software agents (verified 2026-06-12)
- NVIDIA Cosmos 3 physical AI model (verified 2026-06-12)
- NVIDIA physical AI agent tools and skills (verified 2026-06-12)
- NVIDIA Alpamayo 2 Super (verified 2026-06-12)
- NVIDIA RTX Spark Windows AI PCs (verified 2026-06-12)
- NVIDIA DGX Station for Windows (verified 2026-06-12)
- Postman AI Engineer (verified 2026-06-12)
- RelationalAI Snowflake agentic decision intelligence (verified 2026-06-12)
- 7AI Agentic Security Platform (verified 2026-06-12)
- White House AI cybersecurity order (verified 2026-06-12)
Workflow playbooks
- Best Pay-As-You-Go AI Tools and APIs (June 2026)Current buyer guide to true pay-as-you-go AI tools, separating metered APIs from flat subscriptions and showing which platform to use for text, coding, media, voice, and production workloads.
- AI Automation Agency Tech Stack (June 2026)A source-backed AI automation agency stack for selling reliable client workflows without overbuying agent platforms or hiding failure modes.
- Best Open Source AI Tools (June 2026)Current buyer guide to open source and open-weight AI tools, covering local chat, self-hosted interfaces, open models, image generation, speech recognition, privacy tradeoffs, hardware limits, and security risks.
Recent product signals
- GLM-5.2 puts open-model pressure back on closed AI subscriptionsJun 24
- AI News Desk, May 28, 2026: Claude Opus 4.8, Anthropic's $65B round, enterprise agents, wallets, and runtime governanceMay 28
- Compal and GMI Cloud target the infrastructure bottleneck behind large-scale agentic AIMay 28
- AI News Desk, May 27, 2026: OpenRouter funding, Qwen agents, Windows Copilot, and Samsung's multi-model rolloutMay 27
- OpenRouter's $113M Series B makes model routing an enterprise AI infrastructure betMay 27
Spotted an error or want to share your experience with AI Infrastructure & Model APIs?
Every tool page is re-verified on a recurring cycle, and corrections land faster when readers flag them directly. If you spot a stale fact, a missing capability, or have used AI Infrastructure & Model APIs and want to share what worked or didn't, the editorial desk reviews every message sent through this form.
Email editorial@aipedia.wiki