Tencent's WorkForge Turns Real Files Into AI Agent Training Grounds

A new synthesis framework builds 16,700 verifiable training environments across 40 professional domains, lifting a 35B Qwen model's GDPVal score from 45.5 to 73.6.

·
·
·
Tencent's WorkForge Turns Real Files Into AI Agent Training GroundsPRO
  • Researchers from Tencent and Fudan introduce WorkForge, a framework for synthesizing verifiable work-agent training environments.
  • Produces 16.7K environments across 40 professional domains and 60 file types from real-world resources.
  • Workspaces drive task generation, so programmatic and semantic verifiers stay grounded in observable file evidence.
  • Fine-tuned Qwen3.5-35B improves GDPVal from 45.5 to 73.6 and APEX from 5.0 to 21.3.
  • Scaling holds across both training data volume and interaction horizon length.
  • No code or weights released yet; paper is on arXiv as preprint 2610.04906.

WorkForge generates agent-training environments from real files

In a new preprint, researchers from Tencent’s LLM Department and Fudan University introduce WorkForge, a framework that converts web-retrieved files into training environments for AI agents. It maps relationships among the files, extracts checkable facts, generates professional tasks, and builds verifiers that grade the resulting work.

Reliable environments are scarce because professional assignments can span dozens of files, tool calls, and decisions. Post-training requires a workspace in which an agent can act over many steps and receive a defensible score. WorkForge aims to produce those workspaces at data-pipeline scale.

The claim in four numbers

Measure Reported result
Generated environments 16,700
Coverage 40 professional domains and 60 file types
GDPVal, Qwen3.5-35B-A3B-Base 45.5 before fine-tuning, 73.6 after
APEX Score 5.0 before fine-tuning, 21.3 after

Real work strains toy sandboxes

An agent environment packages the material an agent can inspect, the tools it can call, the state that changes during execution, and the tests that score its output. Legal, financial, and research assignments often involve spreadsheets, contracts, presentations, reports, and dependencies that cross file boundaries.

Hand-built environments demand substantial engineering and domain expertise, which limits their scale. Large synthetic collections often simplify the workspace or generate facts with a language model, weakening the connection between a task and its supposed ground truth.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads