Rubicon-preview (30B-A3B)

Rubicon-preview — a 30B-A3B model trained with a novel reinforcement learning framework built on rubric anchors, targeting open-ended, creative and humanities-centric tasks where verifiable rewards are hard to define.

Overview

Rubicon is a reinforcement learning framework that uses rubric anchors to train large language models for open-ended, subjective, and humanities-centric tasks — domains where verifiable rewards are hard to define. The released Rubicon-preview is a 30B-A3B parameter model built on the Qwen3-30B-A3B base.

Unlike conventional RLVR (Reinforcement Learning from Verifiable Rewards), which confines itself to code or math, Rubicon defines fine-grained rubric criteria as training anchors, enabling the model to be optimized for style, emotional expressiveness, and human-like writing.

30B total parameters with 3B activated per token (30B-A3B), trained under the Apache-2.0 license.

Figure 1 — Rubicon's rubric system: offline data filtering feeds rubric instantiation, which drives rubric updating during RL. Source: arXiv:2508.12790.

Highlights

Subjective & open-ended benchmarks (average of 7 tasks)

Model Avg
Qwen3-30B-A3B (base) 65.29
Rubicon-preview 70.50
DeepSeek-V3-671B 68.08
Figure 2 — the creativity vs. instruction-following trade-off. Rubrics that reward strict compliance cost creative/empathic quality, and vice versa. Source: arXiv:2508.12790.

Resources