Tencent Open-Sources AngelSpec: A Unified Training Framework for MTP and Block-Parallel Speculative Decoding on Hy3 Models
Tencent has released AngelSpec, an open-source, torch-native training framework for speculative-decoding draft models. The release covers both autoregressive multi-token prediction (MTP) and the block-parallel DFlash family.
Most speculative-decoding work searches for one drafter that scores well on an averaged benchmark mixture. Real serving traffic does not look like that mixture. AngelSpec treats workload heterogeneity as a first-class design constraint, and specializes structure, training data, and verification depth around it.
Why one universal drafter underperforms
Speculative decoding is lossless. A lightweight drafter proposes several future tokens, and the target model verifies them together in one forward pass using rejection sampling. Acceleration then depends on two things: how many draft tokens get accepted, and how long the complete draft–verify round takes.
Those two quantities move in opposite directions across domains. In high-entropy open-ended conversation, many continuations are semantically valid. The target may pick any one of them, so acceptance decays quickly with proposal depth. Generating and verifying a long block wastes compute. Autoregressive MTP drafting fits this regime because it proposes a shorter candidate sequence.
Code and mathematical reasoning behave differently. Programming syntax, repeated identifiers, formal expressions, and step-by-step derivations constrain future tokens more strongly. These workloads create longer predictable spans, which is exactly what block-parallel drafting amortizes well.
AngelSpec therefore ships two complementary drafters, not one compromise. The MTP model is trained on rich, diverse conversation-oriented data. The block-diffusion model is strengthened with code- and mathematics-focused samples.
The MTP path: Training-Time Test and target-model rollout
The original Hy3 model is trained with a single MTP layer and no recurrent self-conditioned unrolling. At inference the block can be reused recurrently, but the training objective never prepared it for a long self-generated chain. Errors accumulate with depth, so the second and third draft positions accept at substantially lower rates than the first.
AngelSpec addresses this train–inference mismatch with a shared-parameter, multi-depth scheme. It retains D logical prediction depths but reuses one physical MTP block. During training that block is autoregressively unrolled for D steps, with each prediction fed into the next invocation. Following the Training-Time Test principle of EAGLE-3, depth k+1 receives the arg max prediction from depth k instead of the ground-truth token. Parameters are shared, but supervision stays depth-specific: the teacher target advances one future position at every depth.
Two further choices carry most of the gain. First, the target backbone and target language-model head are frozen, and MTP inputs from the backbone are detached. The drafter improves its proposal distribution without touching the distribution that verification must preserve. Second, training uses target-model rollout: responses are generated by the frozen target rather than taken from the original reference corpus. That produces the exact token choices, hidden-state trajectories, and local uncertainty patterns MTP has to approximate at serving time.
The measured effect is concentrated where it should be. At T = 0, mean acceptance moves from 52.8% to 66.4%, and mean accepted length from 2.58 to 2.99. The first-position rate is nearly preserved at 0.799 → 0.814. The deeper positions carry the delta: p3 climbs from 0.290 to 0.706 on GSM8K, and from 0.387 to 0.757 on HumanEval.
DFly: hybrid target conditioning and a predecessor-conditioned AR head
DFly is the block-diffusion architecture. It builds on DFlash with two structural changes.
Hybrid target-conditioning backbone: DFlash concatenates hidden states from multiple target layers and transforms them with a fully connected layer, producing one shared context feature. That single context is then supplied to every draft layer, which limits layer specialization. DFlare instead learns per-layer fusion weights, giving each draft layer its own view of the target hierarchy, but drops DFlash’s learned cross-layer transformation. DFly composes them: the FC branch establishes a common semantic basis, and the per-layer fusion is applied as a residual refinement on top. The extra branch introduces only D × T scalar weights, and its softmax coefficients can be precomputed after training.
Predecessor-conditioned autoregressive head: A parallel backbone predicts each block position from the accepted context only. It cannot see which continuation was actually selected at earlier draft positions, which produces suffix acceptance decay. DFly places a small sequential head after the parallel backbone, converting position-wise marginal predictions into prefix-conditioned distributions. The expensive backbone stays fully parallel; only the small head runs left to right.
Early Performance
On Qwen3-8B, DFly reaches 5.41 average mean accepted length, against 5.32 for DSpark, 4.57 for DFlash, and 3.24 for MTP. It takes the best result on all five math and code benchmarks. DSpark stays slightly ahead on MT-Bench at 3.77 versus 3.67, which is consistent with DFly being positioned for code and math.
On Hy3-A21B, the margin is wider. DFly reaches 4.79 against 3.69 for DFlash and 3.00 for MTP — relative gains of 29.8% and 59.7%. It improves every one of the six reported benchmarks.
The cumulative ablation on Hy3-A21B under greedy decoding traces where that comes from: DFlash backbone 3.77, DFly backbone 4.40, plus Markov head 4.56, and hidden correction 4.60, and code/math data 4.75. The data expansion adds 700K prompts — 500K code from OpenCodeInstruct and OpenCodeReasoning, 200K math from Big-Math. Any prompt sharing a contiguous 16-token span with an evaluation example is removed before response generation.
Inside the framework
AngelSpec is built on TorchSpec and extends it in several places. The foundation is disaggregated: inference engines run the frozen target model and stream hidden states through a Mooncake-backed RDMA store directly to distributed training workers. Hidden states are captured inside vLLM worker processes through public vLLM APIs — a speculative hidden-state extraction hook and a custom KV connector — without forking the engine.
The extensions that matter for this work:
TTT rollout unrolled in parallel over the whole sequence, with the causal-prefix-plus-diagonal attention structure reproduced implicitly via compiled FlexAttention and logsumexp merging. Memory stays close to a single causal pass.
Long-context training with Ulysses sequence parallelism, validated at context lengths up to 128k tokens. Each local shard carries a D-token halo so depth-shifted supervision stays rank-local.
Document-aware sequence packing with three isolation mechanisms — an attention document gate, a depth-shift document gate specific to the MTP path, and document-local position encoding. Cross-document isolation is covered by unit tests verifying zero attention leakage.
Evaluation server that periodically runs genuine speculative decoding against the latest checkpoint on dedicated GPUs, reporting mean accepted length and per-position acceptance as measured by the serving engine itself.
Pluggable interfaces at three levels: targets (runtime vLLM plugin entry point, no source patch), objectives (composed over a shared base, selected by config), and optimizers (Muon as a drop-in alternative to AdamW).
Backends are tiered: vLLM is first-class, with SGLang and HuggingFace Transformers supported at community tier.
Key Takeaways
Seven checkpoints ship on Hugging Face and ModelScope, including no-think and high-think DFly variants.
AngelSpec trains six draft architectures — DFly, DFlash, DFlare, Eagle3, DSpark, MTP — behind one config-driven pipeline.
DFly lifts mean accepted length on Hy3-A21B to 4.79, versus 3.69 for DFlash and 3.00 for MTP.
On HY3-295B-A21B with TP=8, DFly-8 delivers a 1.98–2.40× speedup over autoregressive decoding across concurrency 4 to 64.
D-cut pushes live-traffic throughput to 981 tok/s at c64, +15.7% over DFly, while giving up 2.8% acceptance.
Check out the Paper, GitHub Repo, Documentation, Hugging Face Collection and ModelScope Collection. All credit for this research goes to the researchers of this project.
Michal Sutter is a data science professional with a Master of Science in Data Science from the University of Padova. With a solid foundation in statistical analysis, machine learning, and data engineering, Michal excels at transforming complex datasets into actionable insights.



