#!/usr/bin/env python3 """Build multilingual public README entry points for the project mirrors.""" from __future__ import annotations import json from pathlib import Path ROOT = Path(__file__).resolve().parents[1] UPDATED = "2026-06-23" LANGUAGES = [ ("en", "English", "README.md"), ("zh", "中文", "README.zh.md"), ("es", "Español", "README.es.md"), ("fr", "Français", "README.fr.md"), ("de", "Deutsch", "README.de.md"), ("ja", "日本語", "README.ja.md"), ("ko", "한국어", "README.ko.md"), ("pt", "Português", "README.pt.md"), ] def lang_bar(active: str) -> str: parts = [] for code, label, filename in LANGUAGES: text = f"{label}" if code == active else label parts.append(f' {text}') return "\n
\n" + " ·\n".join(parts) + "\n
\n" def badges() -> str: return """""" def hero(title: str, tagline: str, active: str) -> str: return f"""
{tagline}
{lang_bar(active)} {badges()} """ ENGLISH_TOP = f"""{hero( "Ropedia Xperience-10M Task Suite", "A multilingual public research surface for Xperience-10M: sample data, 20 embodied-AI tasks, baselines, Qwen3-Omni and Cosmos3 diagnostics, and foundation-model training directions.", "en", )} This project builds on the Xperience-10M dataset released by Ropedia to provide public, reproducible embodied-AI evaluation materials. It is organized into two evidence lines. **Line 1** turns one public sample episode into inspectable tasks, targets, and baseline runs. **Line 2** uses selected 128-episode public-safe artifacts for aligned metadata/raw baselines, Qwen3-Omni v6 LoRA, Cosmos3-Super Reasoner, and Cosmos3-Nano Future Window. Every score links back to its source artifact, and direct scores remain clearly separated from compact-proxy estimates. **Updated:** {UPDATED}. **Scope:** Line 1 uses one public sample episode. Line 2 uses selected 128-episode public-safe artifacts linked back to official gated episode paths. Raw Xperience-10M MP4/HDF5/RRD files, Qwen3 base weights, Cosmos3 base weights, and gated data are not redistributed here. ## Contents - [Project Entry Points](#project-entry-points) - [At A Glance](#at-a-glance) - [Data Explorer Analysis](#data-explorer-analysis) - [Two Evidence Lines](#two-evidence-lines) - [Fast Project Map](#fast-project-map) - [Why This Project Exists](#why-this-project-exists) - [Start Here](#start-here) - [Glossary](#glossary) - [Current Research Scope](#current-research-scope) - [Evaluation Protocol](#evaluation-protocol) - [Dataset Context](#dataset-context) - [Reproducibility](#reproducibility) - [Citation](#citation) ## Project Entry Points Use the two evidence lines first, then choose the artifact that answers your question. The dashboard is the best visual overview; the GitHub repo is the source of truth for scripts and generated JSON; Hugging Face mirrors contain public-safe cards, metrics, figures, and model artifacts. Quick rule: use **Line 1** for “can I inspect and reproduce the task?” Use **Line 2** for “how do aligned baselines and model diagnostics compare on the selected 128 episodes?” The multilingual README files provide project overviews. The canonical technical evidence is still the committed task contracts, result matrices, validation JSON, and public-safe result packages. ## At A Glance| Signal | Current public state |
|---|---|
Project identity![]() |
The same project logo mark is used across the GitHub README, GitHub Pages dashboard, Hugging Face Space, artifact dataset, model mirrors, favicon, and social preview. Ropedia is credited as the Xperience-10M data provider and releaser; this repository is the task-suite and evaluation layer built on that dataset. Reusable assets: logo mark and social card. |
| Two-line contract | Line 1: 1 sample episode for task construction and reproducibility. Line 2: 128 selected episodes for same-split metadata/raw baselines, Qwen3-Omni v6, and Cosmos3 diagnostics. |
| Data explorer analysis | A generated analysis layer separates the public sample, selected-128 feature exports, and authenticated Hugging Face gated full-dataset metadata with scope stats, split counts, modality breakdowns, and chart assets. |
| 180 method-task records | 9 methods x 20 tasks = 180/180 scored records. The ledger separates 174 direct scores from 6 compact-proxy scores. |
| 20 task contracts | Action, procedure, transition, trajectory, contact, objects, language, retrieval, reconstruction, order, sync, long-horizon forecasting, interaction text, action-object binding, sensor bridging, camera sync, and transition timing. |
| 4 research directions | Human Modeling & Motion Understanding; 3D/4D Reconstruction & Neural Rendering; Egocentric Vision & Interaction; Scene Reconstruction & World Modeling. These are analysis groups over the same 20 tasks, not separate benchmark tiers. |
| Line 1 methods | Minimal and Neural MLP baselines cover all 20 tasks on the one public sample episode: 40/40 direct scores. |
| Line 2 methods | Metadata simple/NN, raw-feature simple/NN, Qwen3-Omni v6 LoRA, Cosmos3-Super Reasoner, and Cosmos3-Nano Future Window cover all 20 selected-128 task axes: 140/140 scores. |
| 3 foundation pipelines | Spatial intelligence, human-video world modeling, and vision-language-action pipelines are documented as training recipes with task mappings, input-output contracts, and model-evidence requirements. |
| 1 unified target | The long-term embodied foundation-model target connects perception, 3D memory, language-grounded reasoning, action, and planning without adding a new score axis. |
| Public mirrors | GitHub, GitHub Pages, HF Space, HF artifact dataset, HF baseline model repo, Qwen3-Omni and Cosmos3 model repos, and HF collection. |
| Line | Data unit | Score statement | Best use | Read separately from |
|---|---|---|---|---|
| 1 sample episode | One public Xperience-10M sample episode: 5,821 frames, 1,161 aligned 20-frame windows, 8,546 feature dimensions. | 40/40 direct scores from Minimal and Neural MLP heads. | Inspect the raw sample, understand file organization, reproduce the 20 task targets, and compare Minimal vs Neural MLP behavior inside one episode. | The selected-128 comparison rows and any broader held-out model behavior. |
| 128 selected episodes | Selected held-out 96/16/16 split: 34,269 exported windows with public-safe processed features linked to official gated episode paths. The Hugging Face artifact dataset exposes these rows separately as selected_128_windows/selected_128; it is not mixed with the one-sample episode_sample/public_sample viewer. |
140/140 selected-128 scores: 134 direct + 6 compact-proxy. | Compare same-split metadata/raw baselines, Qwen3-Omni v6, Cosmos3-Super, and Cosmos3-Nano while keeping the 6 compact-proxy cells visible. | Direct raw-target measurements for the proxy-marked cells. |
| Line | Methods | Tasks | Scored records | Direct scores | Proxy scores |
|---|---|---|---|---|---|
| 1 sample episode | 2 | 20 | 40/40 | 40 | 0 |
| 128 selected episodes | 7 | 20 | 140/140 | 134 | 6 compact-proxy scores, each source-linked and reasoned. |
| Total public matrix | 9 | 20 | 180/180 | 174 | 6 |
| Evidence line | Method block | Methods | Score statement | Read as |
|---|---|---|---|---|
| 1 sample episode | Task-head baselines | Minimal; Neural MLP | 40/40 direct scores. | Task-lab reproducibility and simple-vs-neural behavior. |
| 128 selected episodes | Aligned baseline heads | Metadata simple/NN; raw-feature simple/NN | 80/80 scores: 74 direct + 6 compact-proxy. | Same-split metadata/raw-feature baseline comparison. |
| 128 selected episodes | Qwen3-Omni series | Qwen3-Omni v6 LoRA | 20/20 direct scores from verified selected-128 Qwen3-Omni LoRA and task-specific probes. | Trainable Qwen3-Omni diagnostic baseline on the selected-128 surface. |
| 128 selected episodes | Cosmos3 series | Cosmos3-Super Reasoner; Cosmos3-Nano Future Window | 40/40 direct scores from verified public-safe reasoner and future-window artifacts. | Cosmos3 reasoner and future-window diagnostics on the selected-128 surface. |
| Run | Purpose | Main change | Eval signal | Use now |
|---|---|---|---|---|
| v1 | Prove the selected-128 LoRA/eval/package loop. | First verified 96/16/16 selected-episode Qwen3-Omni LoRA run. | 448 eval; JSON 0.8750; contact 0.6451. | Lineage only. |
| v2 | Make answers schema-checked. | Structured-JSON contract with full-8-GPU LoRA on the same split. | 448 eval; JSON 0.9978; contact 0.7188. | Structured-output ablation. |
| v3 | Separate prompt/eval effects from training. | Strict-label prompt/eval over the v2 adapter; no new adapter training. | 448 eval; JSON 1.0000; contact 0.7210. | Prompt/eval ablation. |
| v4 | Test longer structured-JSON LoRA training. | New four-epoch full-8-GPU adapter on the same selected split. | 448 eval; JSON 1.0000; contact 0.7299. | Overfit/metric-tradeoff evidence. |
| v5 | Move to denser multiscale evaluation. | Multiscale cap96 export with 4,032 held-out predictions. | 4,032 eval; JSON 1.0000; contact 0.7865. | Pinned prior release; stronger on several non-contact metrics. |
| v6 | Publish the current Qwen 20-task row. | Rank64/lr5e-5 multiscale LoRA plus verified task-specific probes. | 4,032 eval; JSON 0.9990; contact 0.8177. | Current public 20-task Qwen3-Omni row. |
| Goal | Start here | Then inspect |
|---|---|---|
| Understand quickly | Project brief Project status |
Dashboard |
| Choose the published mirror | Public evidence map | evidence-map data |
| Decode project terms | Glossary | glossary data |
| Inspect the 20 tasks | 20-task guide | task contract data task walkthroughs |
| Explore data scales | Data explorer analysis | website analysis section structured analysis record |
| Compare results | Research takeaways | two-line result summary 180-record result table radar data score/proxy audit |
| Understand one sample | Single-episode explorer | sample-file map feature manifest |
| Read foundation directions | Three foundation pipelines | pipeline contract data foundation model plan |
| Reproduce or audit | Reproducibility Evidence contract |
quality gates publication audit mirror parity |