Sentient's EvoSkill Now Runs Self-Improving Coding Agents Without Claude API

EvoSkill v1.3.0 adds Fireworks AI as a first-class inference provider, letting you run fully open-source self-improvement loops at high speed and low cost

·
·
·
Sentient's EvoSkill Now Runs Self-Improving Coding Agents Without Claude API
Read5 min
TypeNews
  • EvoSkill v1.3.0 adds Fireworks AI as a first-class provider for both the agent harness and the LLM scorer.
  • A single FIREWORKS_API_KEY env var is all that's needed; EvoSkill auto-mirrors it across all internal consumers.
  • EvoSkill automatically evolves coding agents by analyzing failures and proposing skill/prompt mutations each iteration.
  • Benchmarks show +7.3% on OfficeQA and +12.1% on SealQA, with zero-shot skill transfer to BrowseComp (+5.3%).
  • Fireworks processes 13T+ tokens/day and leads GPU-based providers on latency, making it well-suited for high-iteration evolution loops.
  • No breaking changes; existing Anthropic, OpenAI, and OpenRouter setups are unaffected. Goose requires manual OpenAI-compatible config.

EvoSkill, Sentient's open-source framework for automatically evolving coding agents, just shipped v1.3.0 with Fireworks AI as a fully supported inference provider. That means you can now run your entire agent self-improvement loop, both the evolution harness and the LLM scorer that judges results, entirely on fast, cheap open-source models. No Claude API required.

What EvoSkill actually does

EvoSkill is an open-source framework that automatically generates structured skills for AI agents by analyzing their failure cases, enabling continuous self-improvement. The core idea is to treat your agent's configuration, its system prompt and learned skills, as something that can be evolved rather than hand-tuned.

The self-improvement loop runs five stages on repeat:

  1. Base Agent attempts benchmark questions using the current best program.
  2. Proposer analyzes failures and suggests targeted skill or prompt changes.
  3. Generator writes the actual new skill files or rewrites the system prompt.
  4. Evaluator scores the new variant on a held-out validation set.
  5. Frontier tracks the top-N performing programs as git branches; the best survive to the next iteration.

EvoSkill significantly extends the feedback-driven idea of GEPA from single-file optimization to complete agent evolution. Instead of only revising one prompt in place, EvoSkill proposes multiple skill and prompt mutations jointly, evaluates new variants on held-out data, and has each iteration produce an entirely new agent program.

The research results back this up. Experiments show consistent improvements on two distinct benchmarks: OfficeQA (+7.3%) and SealQA (+12.1%) using only small training subsets, with direct evidence of zero-shot skill transfer from SealQA to BrowseComp (+5.3%). Skills learned on one task can transfer to completely different ones.

Why Fireworks changes the calculus

Before v1.3.0, running EvoSkill meant paying for Claude or routing through OpenRouter. The new Fireworks integration makes a fully open-source, cost-efficient pipeline possible. Fireworks' engine already runs at internet scale, processing over 13 trillion tokens daily, sustaining about 180 thousand requests per second, and generating over 1,000 tokens per second on large models.

As open models get more powerful and agentic, low latency enables complex, multi-step AI agents to be usable in real-time. Fireworks is fastest across the top open models among GPU-based providers, as benchmarked by Artificial Analysis. For an evolutionary loop that may run dozens of agent episodes per iteration, that throughput matters a lot.

Fireworks gives access to a broad catalog of popular open-source models including DeepSeek, Qwen, Gemma, Kimi, Llama, and Mistral, all optimized for cost, speed, and quality. You can now point EvoSkill at any of these as both the agent harness model and the judge model.

The integration is one environment variable

The v1.3.0 release is deliberately frictionless. Set one key and EvoSkill handles the rest:

code
export FIREWORKS_API_KEY=your-key-here

EvoSkill auto-mirrors FIREWORKS_API_KEY to FIREWORKS_AI_API_KEY for litellm-based harnesses, so you only ever manage one key. The framework also handles automatic provider detection from the model string: fireworks/, fireworks-ai/, fireworks_ai/ prefixes, or a bare Fireworks model ID like accounts/fireworks/models/... all resolve correctly. The key is also forwarded automatically to Docker and Daytona remote runs, so containerized evolution picks it up without extra configuration.

One caveat: Goose has no built-in Fireworks provider, so you will need to configure it manually via an OpenAI-compatible endpoint setup.

To configure Fireworks as both the harness and the LLM scorer in your .evoskill/config.toml:

ini
[harness]
name = "opencode"
model = "fireworks/accounts/fireworks/models/deepseek-v3"
[scorer]
type = "llm"
model = "accounts/fireworks/models/llama-v3p1-70b-instruct"
provider = "fireworks"

What it's good for, and where it falls short

EvoSkill turns general AI agents into state-of-the-art specialists with a benchmark, and is compatible with Claude Code, Codex CLI, OpenCode, OpenHands, Goose, Harbor, and more. The Fireworks backend is especially useful when you need to run many iterations quickly without burning through Claude API credits, or when you want to evolve a specialist agent on top of a fully open-source model stack.

A few things to keep in mind:

  • Goose is excluded from native Fireworks support and needs a manual OpenAI-compatible config.
  • Evolution loops take time. EvoSkill's docs recommend Docker or Daytona for runs that can stretch to hours.
  • Fireworks does one thing well: serve optimized open models fast. But once you need to fine-tune, deploy in your own cloud, or run anything adjacent to the model, it becomes clear what Fireworks is not trying to solve.
  • There is no free tier on Fireworks, though you get $1 in free credits on serverless inference to start.

The bigger picture

The release signals a growing trend in the AI industry toward self-improving systems. Rather than relying solely on human engineers to patch errors, frameworks like EvoSkill aim to create a continuous feedback loop where agents learn from their mistakes in real time.

Adding Fireworks as a first-class provider is a meaningful step toward making that loop accessible and affordable. The combination of fast open-model inference with an automated skill discovery pipeline means you can now run serious agent self-improvement experiments without a large API budget. EvoSkill is available now on GitHub under Apache 2.0, and the Fireworks integration ships with no breaking changes to existing Anthropic, OpenAI, or OpenRouter setups.

Trending
  • No trending articles

Comments

avatar

Next Reads