PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation
Paper • 2609.38597 • Published • 13
The Kling Team is building next-generation multimodal world models across video, audio, text, 3D, and beyond. We are continuously looking for exceptional talent to join us. Feel free to reach out!
ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context Routing
Motion Beyond Morphology: Bootstrapping Cross-Category Motion Transfer from Abstract Motion Representations