ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search
Abstract
ZGCM-1 is a 7B open foundation model that combines internal reasoning with external tool use, trained via efficient architecture-system co-design, progressive long-context scaling, and autonomous agent workflows to achieve strong reasoning and efficiency.
In this work, we present ZGCM-1, a fully open 7B dense foundation model trained from scratch with extreme data, system, and algorithmic efficiency. ZGCM-1 is founded on a core premise: compact models cannot passively memorize the open web, but can overcome parametric capacity limits by coupling deliberate internal thinking with active external tool use. To support this paradigm across a 256K context, we develop an end-to-end, high-efficiency open training recipe: Architecture & System Co-design: interleaved gated sliding-window and full attention, and a stable FP8 Muon optimizer; Progressive Curriculum & MDP Mid-Training: context scaling across 16K, 64K, and 256K, and the reformulation of interaction traces into Markov Decision Processes. Furthermore, we establish an AI-native R&D workflow where agent swarms autonomously manage cluster operations, data curation, and rapid diagnostic evaluation. Extensive evaluations show that ZGCM-1-7B is competitive across 7B model family on general benchmarks. On several challenging mathematical reasoning and agentic search suites, it remains competitive with frontier models orders of magnitude larger, such as Qwen3-235B-A22B and GLM-5.1. We also show that our pre-training design offers a ~4.2x efficiency improvement in 16K pre-training time-to-loss. Across the full development lifecycle, we distill eight actionable empirical findings-spanning architectural scaling, SFT quality pruning, long-context generalization, and agentic co-training dynamics. To facilitate community research, we open-source model weights from the pre-training, mid-training, and post-training stages, intermediate checkpoints, training code, per-stage data and data recipes, and W&B logs.
Community
A fully open 7B LLM, including model, training infra, data, and wandb. Match Qwen3-8B on general datasets and competitive with frontier models orders of magnitude larger, such as Qwen3-235B-A22B and GLM-5.1 on math and agentic search datasets. Its highlight includes FP8 pretraining, MDP midtraining, and AI4AI.
Code: https://github.com/zgcagi/ZGCM-1
Model: https://huggingface.co/zgcagi/ZGCM-1-7B
Data: https://huggingface.co/datasets/zgcagi/ZGCM-1-Data
A big step toward recursive self-improvement (RSI) 🔥🔥🔥
The "AI4AI" part is the real headline — agent swarms running cluster ops and data curation means the model partially built itself. RSI era loading… 🚀
So cool!!!!! Really surprising work! A big step to RSI&AGI.
This is an automated message from the ResearchStudio team.
We created an interactive ResearchStudio Reel for this paper. It includes a visual poster, a video, and a blog, all available for download in editable formats.
Open the ResearchStudio Reel →
Download all files from Hugging Face
Please give this comment a thumbs up if you find the Reel helpful!
Want to explore or create Reels for more papers? Visit the ResearchStudio demo.
This is really cool.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Instella-MoE Technical Report (2026)
- Nanbeige4.2-3B: Unlocking Agentic Capabilities in a Compact Model (2026)
- A.X K2 Technical Report (2026)
- MidTool: Mid-training Data Synthesis for Agentic Tool Use (2026)
- TuringLLM: Efficiently Scaling Foundation Models Toward Physical AI (2026)
- Solar Open 2 Technical Report (2026)
- Kimi K3: Open Frontier Intelligence (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.13356 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash 