Long Chain-of-Thought Compression via Fine-Grained Group Policy Optimization
View original at arxiv.org{ "id": "2602.10048v1", "url": "http://arxiv.org/abs/2602.10048v1", "title": "Long Chain-of-Thought Compression via Fine-Grained Group Policy Optimization", "summary": "Large Language Models (LLMs) often generate unnecessarily verbose Chain-of-Thought (CoT) reasoning that increases computational costs and latency witho…
What we drew from this source
The claims Via News extracted from this document. We point to the source; we don't replace it.
FGO effectively mitigates entropy collapse and preserves sufficient exploration compared to GRPO
80% confidenceReasoning ability does not scale linearly with the length of Chain-of-Thought
80% confidenceFGO consistently achieves 100% data utilization rate across experiments
80% confidenceFGO successfully addresses two major limitations of GRPO: inefficient data utilization and entropy collapse
80% confidenceFGO preserves the majority of self-reflection steps and does not lose reasoning capability despite CoT compression
80% confidenceLarge Language Models often generate unnecessarily verbose Chain-of-Thought reasoning that increases computational costs and latency without proportional performance gains
80% confidenceExcessively long Chain-of-Thought often leads to performance degradation due to overthinking and redundant double-checking
80% confidenceFGO achieves efficient Chain-of-Thought compression without degrading performance
80% confidence
