LoGRA Cuts AI Reinforcement Training Memory by 45% on One Node

A new RL post-training method shrinks memory by 45.7% using low-rank gradient sketches, letting a 27B model train on a single eight-GPU node.

·
·
·
LoGRA Cuts AI Reinforcement Training Memory by 45% on One NodePRO
  • LoGRA compresses RL gradients into low-rank sketches, cutting training memory by up to 45.7% without accuracy loss.
  • Sketches are accumulated during backprop, so the full dense gradient tensor is never materialized.
  • Predicted-KL step control estimates policy drift before each update and shrinks oversized steps to prevent collapse.
  • Trained a 27B model for 1,100+ steps on one 8-GPU node where dense Adam hits OOM.
  • Same sketch doubles as the payload for policy synchronization, reducing inter-worker bandwidth.
  • Code lives in the Molt library with a NeMo-RL pull request in flight.

LoGRA cuts RL post-training memory with low-rank gradient sketches

Reinforcement learning methods such as Group Relative Policy Optimization (GRPO) can improve mathematical reasoning in language models, yet their training loops retain gradients, activations, and optimizer state alongside model weights. The LoGRA paper introduces low-rank gradient sketches for RL post-training. Across the evaluated reasoning tasks, the authors report up to 45.7% lower average training memory and a 27B-parameter model trained for more than 1,100 steps on one eight-GPU node.

Lower memory use can determine whether a training run fits on existing hardware. In the paper’s 27B experiment, dense Adam exhausted available memory under the same tested configuration.

Why Adam fills the GPU

Model weights account for one part of an RL training job’s memory budget. For a 7B-parameter model, the paper’s example breaks down the core allocations as follows:

  • BF16 model weights: approximately 14 GB.
  • Adam moment buffers: approximately 56 GB for two FP32 states.
  • Additional allocations: gradients, activations, rollout data, and framework buffers.

Together, the BF16 weights and Adam moments occupy about 70 GB before the training loop allocates the remaining tensors. The moment buffers alone consume four times as much memory as the BF16 weights.

Reward-based training often derives an update from a small scalar signal attached to each generated response. A binary verifier, for example, may record whether an answer is correct even though backpropagation ordinarily produces a value for every trainable parameter. LoGRA tests the hypothesis that the useful update directions occupy a much smaller space than the dense gradient tensor.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads