Competitive intelligenceEvidence cutoff · 13 августа 2026
Unreal Labs · deep dive

Кто уже стоит
между токеном
и результатом.

Прямые игроки, косвенные владельцы budget line и платформы, которые могут превратить весь слой в бесплатную фичу.

22
ключевых профилей
не logo cemetery
3
слоя конкуренции
direct · adjacent · bundled
3
заметных M&A исхода
Portkey · Langfuse · Helicone
1
surviving wedge
cost per accepted success
12
fresh evidence cards
14 июл → 13 авг · quarantine-reviewed
0
публичных invoice-grade решений
с private rates/credits/corrections
/01
Explore / compare

Три рынка накладываются друг на друга.

Переключай слой и открывай досье. Прямой конкурент — не тот, у кого есть cost dashboard, а тот, кто занимает тот же control point или способен связать trace, eval и изменение системы.

Прямые · Evaluation + observability

Arize AI / Phoenix

Independent · scaled · Founded 2020; Phoenix repo since Nov 2022

Высокий overlap

Cross-framework observability/evaluation for ML, LLMs and agents; Phoenix OSS; AX includes guardrails, prompt learning and managed agent workflows.

Кто покупает
ML/AI platform and reliability teams
Business model
Open-source/source-available distribution + AX SaaS: Free, Pro $50/mo, enterprise custom.
Капитал / сделки
$70M Series C in February 2025; ~$131M total raised.
Revenue
Undisclosed.
Traction
11k Phoenix GitHub stars on 2026-08-13; company reported 2M+ monthly downloads at Series C.
Инвесторы
Adams Street Partners · M12 · Sinewave · OMERS Ventures · Datadog · PagerDuty · Foundation Capital · Battery · TCV
Где пересекается
evaluationobservabilityruntime guardrailsprompt learning
Не закрывает

Capabilities exist, but public independently measured savings and invoice-grade workflow economics do not.

Следить

Сильный counterexample против тезиса 'никто не замыкает eval на intervention'.

Long tail

Игроки, которых нужно мониторить, но не обязательно превращать в одинаковые досье.

Gateways

nexos.ai · Bifrost · Envoy AI Gateway · A10 AI Gateway · Apigee AI · Netlify AI Gateway · Vercel AI Gateway

Observability / evals

W&B Weave · AgentOps · Maxim AI · Keywords AI · Traceloop/OpenLLMetry · New Relic AI Monitoring · Raindrop

FinOps / spend

IBM Apptio/Cloudability · Kubecost · Harness CCM · PointFive · nOps · Pay-i · Ramp Token Spend · Amnic

Routers / optimization

Unify · RouteLLM implementations · AWS Intelligent Prompt Routing · NVIDIA NeMo Switchyard · OpenAI/Anthropic native model selection

Fresh watch · last 30 days

BurnLens · TokenSpend · TokenMaxxer · Wattage · TokenShield · Aurora · A10 AI Gateway

/02
last30days / discovery layer

Что изменилось за последние 30 дней.

Не новостная лента и не автоматический импорт claims. Это quarantine-reviewed слой: свежие инциденты, launches, operator language и продуктовые формы, которые могут сдвинуть карту между большими research-волнами.

14 июля — 13 августа 202613 августа 2026 · 17:30 EEST

4 узких last30days-прохода: incidents, reconciliation, products/bundling и unit economics. Reddit был частичным из-за 429; X дополнен через официальный xurl. Browser cookies не использовались.

Collection health
Hacker News · ok
launches + operator discussion
GitHub · ok
issues, repos, maintainer evidence
YouTube · ok
transcripts; vendor talks separated
Reddit · partial
RSS/archives; HTTP 429
X · supplemented
read-only xurl pass
Failure amplification

«Busy» не значит «progress»

Повторные model calls могут постоянно обновлять heartbeat, пока workflow часами не производит результата. Нужны semantic-progress и spend-rate circuit breakers, а не только timeout по тишине.

Metric convergence

Cost per completed task

В r/FinOps и новых продуктах появился тот же denominator: failed/rejected attempts и rework должны быть отнесены к accepted outcomes. Это совпадение языка — сильный discovery signal, но ещё не buyer proof.

Platform bundling

Gateway становится baseline

Google, Microsoft и A10 объединяют routing, quotas, telemetry, security и model abstraction. Standalone gateway commoditizes быстрее, чем успевает оформиться отдельная category.

Nascent competition

Wedge уже копируют по частям

BurnLens, TokenSpend, Wattage и TokenShield закрывают accepted-outcome attribution, PR linkage, regression gates и circuit breakers. Почти нулевой traction не отменяет product convergence.

incident
3 августа 2026

1 333 retry-вызова и $204.74 без forward progress

Два OpenClaw-инцидента длились 3+ часа каждый: activity clock обновлялся на каждом model retry, поэтому stall detector не срабатывал. Maintainer подтвердил дефект и merged fix на main.

Почему важно · primary incident

Редкий публичный артефакт, где failure → calls → wall clock → metered cost связаны в одну цепочку.

Высокая для инцидента

Один проект и одна конфигурация fallback; не prevalence и не enterprise WTP. Fix ещё не был в tagged release на момент проверки.

Source · GitHub issue + maintainer reproduction
practice
7 августа 2026

В scale-практике hard caps — крайний случай, routing — основной lever

Databricks описывает model/harness flexibility, eval-backed routing, visibility, tripwires и progressive gates; internal Smart Router заявлен как >30% average task-cost reduction при близком качестве.

Почему важно · vendor practice

Сильный buyer/operator framing от раннего large-scale adopter; HN discussion: 315 points, 268 comments.

Средняя

Результат self-reported; savings основаны на internal data и informal survey, не независимом benchmark.

Source · Databricks engineering post + HN
launch
13 августа 2026

A10 выпустила AI Gateway как единый control plane

Routing, cost management, governance и security собираются в enterprise network product, а не отдельную token-cost категорию.

Почему важно · official launch

Ещё один incumbent превращает gateway/control layer в bundled capability.

Высокая для capability

Launch claim не показывает adoption, ROI или superiority.

Source · A10 official launch
platform
4 августа 2026

Google Cloud унифицирует model routing на API Gateway

Provider abstraction и model routing перемещаются в API-platform layer — туда, где уже живут policy, auth и enterprise traffic.

Почему важно · official product

Подтверждает commoditization standalone gateway и давление платформенной дистрибуции.

Высокая для capability

Capability announcement; не доказательство оптимальной маршрутизации или savings.

Source · Google Developers Blog
platform
22 июля 2026

Azure APIM бандлит token limits, metrics, cache, fallback и agent/tool governance

Microsoft публично позиционирует existing API management как AI gateway для моделей, MCP tools и A2A agents.

Почему важно · vendor talk

Показывает, что enterprise buyer может купить control plane внутри уже принятого API stack.

Высокая для capability

Vendor conference talk; не независимая оценка использования или качества.

Source · Microsoft PM talk · INTEGRATE 2026
product
проверено 13 августа 2026

BurnLens прямо заходит в cost per accepted outcome

Open-source proxy обещает hard caps до вызова, attribution по workflow/agent и denominator «total workflow spend / accepted outcomes», включая failed/rejected attempts и rework.

Почему важно · nascent product

Самый прямой свежий overlap с surviving wedge — формулировка почти совпадает.

Высокая для surface · низкая для traction

3 GitHub stars, 0 forks; outcome logic и $5.03/merged PR self-reported на собственном repo. Нет buyer/retention evidence.

Source · Product docs + GitHub
product
9 августа 2026

TokenSpend связывает coding spend с shipped PR

Продукт заявляет team-level spend, автоматический PR linkage и shipped-rate вместо per-engineer leaderboard.

Почему важно · nascent product

Сигнал, что category vocabulary сдвигается от tokens к тому, что бюджет реально shipped.

Высокая для surface · низкая для traction

Show HN: 3 points, 2 comments; claims не подтверждены customer artifacts. Название «AI ROI» шире фактически показанной coding/PR surface.

Source · Product page + Show HN
product
26 июля 2026

Wattage превращает token efficiency в CI regression gate

OTel trace profiler ищет context churn, retries, model mismatch и non-convergence; risky interventions требуют quality evidence до зачёта в score.

Почему важно · open-source launch

Очень близок к structural diagnosis + eval-gated intervention, но на dev/CI surface.

Высокая для implementation · низкая для traction

21 GitHub star; 10-loop benchmark synthetic, production efficacy и buyer demand не подтверждены.

Source · GitHub + Show HN
language
11 августа 2026

Operator language сходится на cost per completed task

Пост формулирует метрику как cost per attempt / success rate + стоимость escaped wrong outputs — почти ровно наш fully loaded denominator.

Почему важно · community signal

Ценный язык для интервью и landing copy; указывает, что raw per-token price перестаёт быть основной decision unit.

Низкая

1 point, 1 comment; возможно vendor seeding. Нельзя считать независимым подтверждением спроса.

Source · r/FinOps
incident
18 июля 2026

$58 за validation → retry → huge JSON loop

Автор описывает агента, который не crash-нулся, а продолжал генерировать невалидный output и повторять попытки несколько часов.

Почему важно · community anecdote

Живой operator-language пример «failure that looks active» и stochastic tax.

Низкая

Анонимный рассказ, 11 comments, без bill artifact; только hypothesis-generating evidence.

Source · r/LocalLLM anecdote
benchmark
12 августа 2026

Одинаковая accuracy, 6× разный cost per task

В CursorBench 3.2 обсуждались 70.8% при $2.81/task против 70.5% при $17.32/task — иллюстрация, почему модель выбирают по frontier quality/cost, а не benchmark rank.

Почему важно · social benchmark

Хорошая свежая визуализация eval-gated routing thesis.

Средняя для quoted benchmark

Benchmark ≠ production workflow; harness, task mix и pricing критичны. Source — X clip, не полноценный methodology paper.

Source · X video clip / CursorBench discussion
product
26 июля 2026

TokenShield продаёт circuit breaker против loop/token bleed

Новый lightweight proxy появился ровно вокруг runaway tool loops, context ballooning и pre-spend enforcement.

Почему важно · self-reported launch

Подтверждает emerging product shape вокруг failure amplification.

Низкая

Self-promotional post, 1 point, 0 comments; adoption и reliability неизвестны.

Source · r/LlamaIndex launch post
/03
What the map says

Четыре давления формируют нишу.

Эта карта важна не количеством логотипов, а тем, какое стратегическое пространство остаётся после consolidation, bundling и convergence.

Consolidation
3 exits

Portkey → Palo Alto; Langfuse → ClickHouse; Helicone → Mintlify. Возможные чтения противоположны: strategic validation или неспособность pure-play расти самостоятельно.

Commoditization
$0 gateway

Cloudflare и Vercel могут отдавать routing/telemetry почти бесплатно, потому что монетизируют cloud, credits и deployment.

Outcome gap
request ≠ success

Span/token/request единицы насыщены. Wedge начинается там, где failure, retry, human review и accepted result меняют denominator.

Accounting gap
invoice ≠ estimate

Private rates, commitments, credits и corrections плохо соединяются с application traces. Это наиболее устойчивый residual gap.

Surviving strategic wedge
Не ещё один gateway. Не ещё один dashboard.

Invoice-accurate attribution → accepted successful workflow → structural diagnosis → customer-approved, eval-verified intervention.

Это не greenfield. Возможность существует только как более узкий outcome/control wedge поверх или внутри существующих платформ.

/04
Reading room

Не список ссылок. Маршрут погружения.

Papers объясняют границы routing и agents; reports — рынок и buyer context; engineering posts — реальные рычаги; social/HN — где практики спорят и где появляются новые category narratives. Fresh items из last30days уже снабжены evidence boundary и не смешаны с verified financial claims.

01
paper
routing

FrugalGPT: How to Use Large Language Models While Reducing Cost

Chen, Zaharia, Zou · 2023

LLM cascades могут резко снижать cost на отдельных датасетах.

Зачем читать

Исток headline 98% savings — и хороший урок, почему максимум на одном dataset нельзя продавать как typical production result.

Caveat

Диапазон зависит от dataset; 98% — максимум, не универсальный эффект.

02
paper
routing

RouteLLM: Learning to Route LLMs with Preference Data

Ong et al. · 2024/ICLR 2025

Router может удерживать 95% качества сильной модели с меньшим числом дорогих вызовов.

Зачем читать

Лучший мост от routing hype к distribution shift и OOD failure.

Caveat

На MMLU router, обученный на Arena, был близок к random; экономия требует task-specific evals.

03
paper
routing

RouterBench: A Benchmark for Multi-LLM Routing Systems

Hu et al. · 2024

Несколько умных routers не превосходят простой Zero Router на части задач.

Зачем читать

Антидот против идеи, что model routing автоматически становится moat.

Caveat

Benchmark не покрывает все production distributions и enterprise constraints.

04
paper
routing

Rerouting LLM Routers

research paper · 2025

Adversarial confounders заставляют routers чаще выбирать дорогие модели.

Зачем читать

Показывает, что optimization control surface сам становится attack surface.

Caveat

Adversarial setup; не incidence rate в обычном production traffic.

05
paper
agents

Beyond Function Calling: Tool-Using Agents under Environment Unreliability

research paper · 2026

После failure 44–76% trajectories повторяют тот же tool.

Зачем читать

Прямое evidence, почему request cost занижает cost of successful workflow.

Caveat

Benchmark behavior, не aggregate production retry rate.

06
paper
context

Lost in the Middle

Liu et al. · 2023

Больше context не гарантирует лучшее качество; позиция информации важна.

Зачем читать

Связывает token reduction с quality: waste нельзя определять только объёмом.

Caveat

Модели и context handling эволюционируют; эффект надо переоценивать на текущем stack.

07
report
market

State of FinOps 2026

FinOps Foundation · 2026

98% respondent FinOps practices управляют AI spend; granular AI monitoring — top request.

Зачем читать

Сильнейший сигнал, что проблема входит в существующую дисциплину.

Caveat

Это выборка FinOps respondents, не 98% всех предприятий; не доказывает standalone WTP.

08
report
market

2025 State of Generative AI in the Enterprise

Menlo Ventures · 2025

Menlo оценивает enterprise GenAI spend в $37B.

Зачем читать

Размер spend pool и application layer context.

Caveat

Survey/modelled estimate; Menlo restated prior-year denominator and excludes/changes some categories.

09
report
usage

State of AI: 100 Trillion Token Study

OpenRouter · 2026

Programming и reasoning резко выросли в token mix; рынок моделей фрагментируется.

Зачем читать

Редкий large-scale usage dataset.

Caveat

OpenRouter traffic explicitly excludes enterprise/private/ZDR usage; task classification sampled.

10
report
observability

State of AI Engineering

Datadog · 2026

69% input tokens в sample были system prompts; cached reads использовались лишь в 28% calls.

Зачем читать

Дает production-shaped optimization hypotheses.

Caveat

Datadog customer telemetry, не весь рынок; low cache rate не доказывает waste автоматически.

11
engineering
agents

Expensively Quadratic: The LLM Agent Cost Curve

exe.dev · 2026

Replay растущей истории делает agent loops дороже нелинейно.

Зачем читать

Лучшее интуитивное объяснение скрытого fully loaded cost длинной trajectory.

Caveat

Авторская модель и конкретный harness; не универсальный benchmark.

12
HN thread
agents

HN discussion: Expensively Quadratic

Hacker News community · 2026

Практики спорят о cache reads, tool outputs, context trimming и observability.

Зачем читать

Полезная peer review поверх engineering post.

Caveat

Community anecdotes; upvotes не evidence of buyer demand.

13
engineering
context

Code execution with MCP: building more efficient AI agents

Anthropic · 2025

Один workflow снизился примерно с 150k до 2k токенов.

Зачем читать

Показывает, что architectural change может дать orders-of-magnitude больше, чем price shopping.

Caveat

Vendor-selected example; не portfolio-wide savings distribution.

14
engineering
context

Code Mode: an entire API in 1,000 tokens

Cloudflare · 2026

2,500+ endpoints: >1.17M tool-definition tokens → ~1k.

Зачем читать

Очень сильный пример structural token reduction.

Caveat

Input schema tokens, не end-to-end cost per successful task.

15
benchmark
agents

MCP vs CLI: Agent Cost & Reliability

Scalekit · 2026

В 75-run benchmark MCP тратил 4–32x tokens и был менее надёжен, чем CLI.

Зачем читать

Редкий simultaneous cost+reliability comparison на tool workflows.

Caveat

Один setup: Claude Sonnet 4, GitHub tasks и конкретный MCP/CLI design.

16
engineering
caching

How We Cut LLM Costs by 59% With Prompt Caching

ProjectDiscovery · 2026

Cache hit rate вырос 7%→84%, vendor reports 59% cost reduction.

Зачем читать

Production-shaped before/after на 20–40+ step agent.

Caveat

Self-reported case; quality and full workload denominator need scrutiny.

17
report
economics

Artificial Analysis methodology

Artificial Analysis · 2026

Cost per task учитывает input, cache read/write, reasoning и answer tokens.

Зачем читать

Нужный denominator shift от price/token к task economics.

Caveat

Benchmark task ≠ accepted customer workflow; human/retry/infra costs may be outside.

18
X thread
platform

AI tokens are another software-engineering resource

Matei Zaharia · 2026-08-07

Databricks frames AI Gateway as centralized analysis/control for token usage.

Зачем читать

Знаковый incumbent signal: control plane идет в platform engineering.

Caveat

Product narrative from vendor founder, not independent market proof.

19
X thread
observability

Stop paying the judge tax

Ben Hylak / Raindrop · 2026-08-11

Rare failures <1% can disappear under sampled observability; LLM judges create their own cost.

Зачем читать

Показывает second-order economics evaluation itself.

Caveat

Competitive product claim; 20B traces/month is vendor-reported.

20
X post
routing

NVIDIA NeMo Switchyard launch

NVIDIA · 2026-08-11

Routing each workflow step is becoming a platform capability.

Зачем читать

Strong bundling signal from the infrastructure layer.

Caveat

Launch announcement and customer claims, not neutral benchmark.

21
X post
agents

A Codex project would have cost $23.28 at API prices

Simon Willison · 2026-08-07

Subscription economics can hide the true marginal API-equivalent cost.

Зачем читать

Useful micro-example of cost attribution across pricing models.

Caveat

One personal project; API-equivalent estimate, not realized invoice.

22
HN thread
economics

Why current LLM costs are not sustainable

Aditya Patadia + HN · 2026

Debate spans open weights, subscription subsidy, local models and falling unit cost.

Зачем читать

Strong counter-case to a durable token-price pain thesis.

Caveat

Opinion essay and community debate; several forward-looking claims are speculative.

23
engineering
finops

Managing AI Coding Costs at Scale

Databricks · 2026

Routing, spend gates, caching, context reduction, evals and gateway converge into one platform problem.

Зачем читать

Almost the complete strategic threat map in one incumbent post.

Caveat

Vendor architecture narrative, not proof of customer WTP for every component.

24
case study
finops

AI Cost Optimization Across 50+ LLMs

CloudZero · 2025/updated 2026

Anonymous SaaS case reports $1M+ savings and attribution across model/customer segments.

Зачем читать

Rare production-shaped FinOps outcome case.

Caveat

Vendor-hosted, customer unnamed, no independent before/after audit.

25
engineering
failure economics

Runaway model-call retry loop bills $204 over two incidents

OpenClaw issue + maintainer review · 3 Aug 2026

Activity-based liveness missed a busy retry stall: 1,333 calls, no final reply, $204.74.

Зачем читать

Лучший свежий публичный incident artifact для semantic progress, retry ceilings и circuit breakers.

Caveat

Один OSS project; не prevalence. Fix landed on main after the incident.

26
engineering
platform bundling

A unified API for AI model routing

Google Developers Blog · 4 Aug 2026

Model abstraction и routing встраиваются в cloud API gateway.

Зачем читать

Показывает, почему standalone routing layer испытывает bundling pressure.

Caveat

Official capability post, не comparative benchmark.

27
engineering
platform bundling

Governing Models, Tools, and Agents with Azure API Management

Microsoft · INTEGRATE 2026 · 22 Jul 2026

Token limits, metrics, routing, fallback и MCP/A2A governance становятся частью обычного API management.

Зачем читать

Полный transcript-level обзор bundled enterprise substitute.

Caveat

Vendor talk; adoption и ROI не раскрыты.

28
engineering
outcome economics

BurnLens: cost per accepted outcome

Sairin Technology · checked 13 Aug 2026

Hard cap до вызова + workflow attribution + denominator с failed/rejected attempts.

Зачем читать

Самый прямой свежий overlap с нашим wedge — полезен для product teardown.

Caveat

3 stars; self-reported metrics на собственном repository.

29
engineering
cost regression

Wattage: token-spend profiler and cost-regression gate

Faizan Raza · 26 Jul 2026

Trace detectors и CI gates связывают structural waste с quality-risk evidence.

Зачем читать

Практическая реализация diagnosis → intervention guardrail.

Caveat

21 stars; convergence benchmark synthetic.

30
social
operator language

Cost per task as the unit metric for AI spend

r/FinOps · 11 Aug 2026

Failed attempts и escaped errors должны входить в стоимость завершённой задачи.

Зачем читать

Язык для interview guide и product framing.

Caveat

1 point / 1 comment; возможно vendor seeding.

31
social
benchmark

CursorBench 3.2: similar accuracy, very different cost per task

MTS / Cursor discussion on X · 12 Aug 2026

70.8% при $2.81/task против 70.5% при $17.32/task в quoted benchmark.

Зачем читать

Свежий пример качества как constraint для routing.

Caveat

Benchmark clip, не production TCO; methodology нужно читать отдельно.

32
social
new products

TokenSpend: AI coding spend → shipped PR

Show HN / TokenSpend · 9 Aug 2026

Attribution сдвигается к shipped code и team-level spend.

Зачем читать

Nascent category signal и direct product overlap.

Caveat

3 HN points, 2 comments; traction не доказан.

Competitive landscape · 13 августа 2026

Fresh ≠ proven: community posts, launch pages, stars и engagement показывают язык, инциденты и emerging products, но не revenue, retention, enterprise adoption, prevalence или WTP.

Company dossiers опираются на product docs, официальные funding/M&A announcements, GitHub metadata и claim ledger. Отдельный слой последних 30 дней собран last30days по Reddit, HN, YouTube и GitHub, дополнен X через read-only xurl и прошёл ручной quarantine review.