Factory Acquires Mentlio to Cut Coding Agent Costs by 30%

Factory folds in YC-backed Mentlio to bring intelligent model routing and token-level ROI measurement into its autonomous software engineering platform.

·
·
·
Factory Acquires Mentlio to Cut Coding Agent Costs by 30%
Read6 min
TypeNews
SubtopicCode Agents
  • Factory acquired Mentlio (YC S26), a token optimization and engineering intelligence startup.
  • Mentlio's two founders, Ashank Shah and Ahmet Demirbas, join Factory; deal terms undisclosed.
  • Core tech: in-harness model routing, context compression, and per-PR ROI measurement.
  • Mentlio claims ~30% AI spend reduction and 98.93% of frontier accuracy at 24.9% lower cost.
  • Follows Factory's $200M raise at $5B valuation and Menlo Ventures investment.
  • Signals coding agent vendors absorbing FinOps and observability layers into their core platforms.

Factory brings Mentlio in-house to track coding-agent economics

Factory has acquired Mentlio, a two-founder Y Combinator startup that measures coding-agent costs, compresses context, and routes prompts to lower-cost models when quality permits. Factory announced the acquisition without disclosing its price, structure, integration schedule, or effect on Mentlio’s standalone product.

Engineering teams can already see how many tokens an agent consumes. Connecting that bill to accepted code, developer productivity, and business outcomes remains harder. Mentlio gives Factory technology for measuring those relationships while reducing the cost of individual agent runs.

What Factory bought, and what remains unclear

Question Current answer
Who is joining Factory? Mentlio founders Ashank Shah and Ahmet Demirbas.
What did Mentlio build? Model routing, context compression, token-cost analysis, and engineering output metrics.
What were the deal terms? Factory did not disclose the price, equity, or transaction structure.
When will Factory ship the technology? No integration schedule or general-availability date was announced.
Will Mentlio remain available? Factory did not describe plans for the standalone product or its customers.

Two founders focused on one expensive problem

Mentlio was founded by Caltech computer science graduates Shah and Demirbas and joined Y Combinator’s Summer 2026 batch. According to Factory, the pair met at Caltech, worked part-time at startups, and won first place at Carnegie Mellon University’s NexHacks with a token-compression algorithm that helped inspire the company.

The small team makes the transaction a founder-and-technology acquisition rather than a conventional product rollup. Factory gains engineers who have worked on model selection and context efficiency, two cost controls that become more valuable as agents take on longer tasks.

Factory, founded by Matan Grinberg and Eno Reyes, develops the Droid software-engineering agent. The company recently announced a $200 million funding round at a $5 billion valuation and operates offices in San Francisco, New York, London, and Japan.

How Mentlio cuts an agent’s bill

Mentlio combined an engineering intelligence dashboard with several mechanisms for reducing model usage. Its materials claim savings of about 30%, although Factory has yet to publish production results from a Droid integration.

  • ROI measurement: Mentlio converts pull requests into normalized “Delivery Points,” giving finance and engineering teams a proposed measure of cost per unit of output.
  • Intelligent routing: The system selects a lower-cost model when it predicts that model can complete the prompt within the required quality threshold.
  • Cache-aware decisions: The router accounts for provider discounts on cached input. Switching models can discard those savings, so a cheaper model does not always produce a cheaper request.
  • Context compression: Mentlio removes less relevant material from long prompts and conversation histories while attempting to preserve task quality.
  • Cost analysis: Dashboards break down token use by engineer, team, tool, and workflow.

The routing layer evaluates prompts inside the agent’s execution loop. That position gives it access to the task, available context, tool results, and prior model calls. A generic API proxy usually sees less information and may struggle to distinguish a routine file lookup from a repository-wide refactor.

Privacy claims need implementation detail

Mentlio says its software computes workflow, quality, and productivity signals on the client side with proprietary machine-learning models. Its dashboards receive aggregate scores rather than prompts, source code, or private files. That design could reduce the amount of sensitive engineering data sent to an additional cloud service.

Enterprise reviews will still require details about what runs locally, which metadata leaves the customer environment, how identifiers are aggregated, how long telemetry is retained, and whether any data is used for model training. Factory has not yet described how Mentlio’s architecture will fit Droid’s deployment and security controls.

A precise benchmark with open variables

Mentlio reports that one router achieved 98.93% of the accuracy of its Fable baseline at 24.90% lower cost on Terminal-Bench 2.1, a benchmark for agents performing terminal-based tasks. The figures come from Mentlio and should be read as a vendor benchmark until broader evaluations are available.

Useful validation would identify the tested models, prompt distribution, cache state, provider pricing, latency, retry rates, run variance, and statistical confidence. Repository language and task type also matter because a router calibrated on shell tasks may behave differently on migrations, debugging sessions, or security-sensitive changes.

Factory needs better unit economics

Factory lists NVIDIA, Blackstone, Palo Alto Networks, Adyen, and EY among its customers. Large deployments expose the economics of autonomous agents quickly because a single task can trigger repeated planning, code generation, tool use, testing, review, and repair calls.

An agent that sends every step to the most capable and expensive model can raise customer bills or compress the vendor’s margins. Per-prompt routing offers a way to reserve frontier models for difficult reasoning while assigning routine subtasks to cheaper models. Context compression can lower input charges further, especially during long sessions that repeatedly send repository history.

Placing those controls inside Droid could also improve observability. Factory would be able to record which model handled each step, why the router selected it, how much the step cost, and whether the resulting change passed tests or review. Developers will need access to that trace when diagnosing regressions caused by routing decisions.

Competition extends beyond agent quality

Cursor, Cognition, Claude Code, and other coding-agent providers face similar cost and routing decisions. As model capabilities converge on common software tasks, vendors can compete through model orchestration, context management, latency, reliability, and cost per accepted change.

Factory’s recent financing gives it capital to build those systems and expand enterprise distribution. The acquisition also follows Chris Degnan’s move from a Factory board-adviser role to become Cognition’s chief revenue officer. Grinberg publicly alleged that Degnan might have shared confidential information with Cognition; the supplied announcement established no connection between that dispute and the Mentlio transaction.

What developers should ask before rollout

Developers evaluating a future Mentlio-powered Droid release will need concrete answers about routing behavior and measurement:

  • Which models, languages, repositories, and task types does the router support?
  • Can teams pin a model for security-sensitive or high-risk work?
  • What confidence threshold triggers a cheaper model, and who controls it?
  • How does the system recover when the selected model fails or produces low-quality code?
  • Do reported costs include cache discounts, retries, tool calls, and failed runs?
  • Can developers inspect routing decisions and compare them with a fixed-model baseline?
  • How are Delivery Points calibrated across teams, and what prevents metric gaming?
  • What happens to existing Mentlio customers, data, and integrations?

Attribution remains the harder test

Standalone AI observability and LLM cost-management vendors will face more pressure if agent platforms bundle routing, compression, and ROI reporting into their execution environments. Model providers face mixed effects: smarter routing can reduce expensive calls per task while lower operating costs may encourage broader agent deployment.

Delivery Points could help teams compare spending across workflows, but a proprietary score requires validation against durable outcomes. A merged pull request remains an intermediate output; reliability, security, cycle time, maintenance burden, and revenue impact determine whether the work created business value.

Factory now has a team focused on linking model spend to engineering output. Proof will require a published integration, reproducible evaluations, stable quality and latency, auditable data handling, and evidence that its output metrics correspond to results customers already track.

Trending
  • No trending articles

Comments

avatar

Next Reads