-
Warp-as-History: Generalizable Camera-Controlled Video Generation from One Training Video
Paper • 2605.15182 • Published • 39 -
STALE: Can LLM Agents Know When Their Memories Are No Longer Valid?
Paper • 2605.06527 • Published • 47 -
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis
Paper • 2605.14392 • Published • 10 -
World Action Models: The Next Frontier in Embodied AI
Paper • 2605.12090 • Published • 71
Collections
Discover the best community collections!
Collections including paper arxiv:2605.26089
-
YOLO-Master: MOE-Accelerated with Specialized Transformers for Enhanced Real-time Detection
Paper • 2512.23273 • Published • 15 -
A 58-Addition, Rank-23 Scheme for General 3x3 Matrix Multiplication
Paper • 2512.21980 • Published • 3 -
Step-DeepResearch Technical Report
Paper • 2512.20491 • Published • 89 -
SAM Audio: Segment Anything in Audio
Paper • 2512.18099 • Published • 25
-
FAN: Fourier Analysis Networks
Paper • 2410.02675 • Published • 29 -
Tensor Product Attention Is All You Need
Paper • 2501.06425 • Published • 91 -
Scalable-Softmax Is Superior for Attention
Paper • 2501.19399 • Published • 25 -
EQ-VAE: Equivariance Regularized Latent Space for Improved Generative Image Modeling
Paper • 2502.09509 • Published • 9
-
Why Fine-Tuning Encourages Hallucinations and How to Fix It
Paper • 2604.15574 • Published • 26 -
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
Paper • 2604.24763 • Published • 71 -
Programming with Data: Test-Driven Data Engineering for Self-Improving LLMs from Raw Corpora
Paper • 2604.24819 • Published • 91 -
GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents
Paper • 2604.26752 • Published • 114
-
Seedream 4.0: Toward Next-generation Multimodal Image Generation
Paper • 2509.20427 • Published • 88 -
Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer
Paper • 2511.22699 • Published • 249 -
RealGen: Photorealistic Text-to-Image Generation via Detector-Guided Rewards
Paper • 2512.00473 • Published • 27 -
Diversity-Preserved Distribution Matching Distillation for Fast Visual Synthesis
Paper • 2602.03139 • Published • 45
-
Warp-as-History: Generalizable Camera-Controlled Video Generation from One Training Video
Paper • 2605.15182 • Published • 39 -
STALE: Can LLM Agents Know When Their Memories Are No Longer Valid?
Paper • 2605.06527 • Published • 47 -
Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesis
Paper • 2605.14392 • Published • 10 -
World Action Models: The Next Frontier in Embodied AI
Paper • 2605.12090 • Published • 71
-
Why Fine-Tuning Encourages Hallucinations and How to Fix It
Paper • 2604.15574 • Published • 26 -
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
Paper • 2604.24763 • Published • 71 -
Programming with Data: Test-Driven Data Engineering for Self-Improving LLMs from Raw Corpora
Paper • 2604.24819 • Published • 91 -
GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents
Paper • 2604.26752 • Published • 114
-
YOLO-Master: MOE-Accelerated with Specialized Transformers for Enhanced Real-time Detection
Paper • 2512.23273 • Published • 15 -
A 58-Addition, Rank-23 Scheme for General 3x3 Matrix Multiplication
Paper • 2512.21980 • Published • 3 -
Step-DeepResearch Technical Report
Paper • 2512.20491 • Published • 89 -
SAM Audio: Segment Anything in Audio
Paper • 2512.18099 • Published • 25
-
Seedream 4.0: Toward Next-generation Multimodal Image Generation
Paper • 2509.20427 • Published • 88 -
Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer
Paper • 2511.22699 • Published • 249 -
RealGen: Photorealistic Text-to-Image Generation via Detector-Guided Rewards
Paper • 2512.00473 • Published • 27 -
Diversity-Preserved Distribution Matching Distillation for Fast Visual Synthesis
Paper • 2602.03139 • Published • 45
-
FAN: Fourier Analysis Networks
Paper • 2410.02675 • Published • 29 -
Tensor Product Attention Is All You Need
Paper • 2501.06425 • Published • 91 -
Scalable-Softmax Is Superior for Attention
Paper • 2501.19399 • Published • 25 -
EQ-VAE: Equivariance Regularized Latent Space for Improved Generative Image Modeling
Paper • 2502.09509 • Published • 9