Shanghai Innovation Institute (SII) · Fudan University · MOSI.AI
Open research on foundation models for language, perception, speech, and embodied intelligence.
OpenMOSS is led by Prof. Xipeng Qiu at the Shanghai Innovation Institute (SII), in collaboration with Fudan University and MOSI.AI. We release models, datasets, benchmarks, and research tools for language, multimodal perception, speech, and embodied intelligence.
| Direction | Flagship repositories |
|---|---|
| 🧠 Language models | MOSS · DiRL · BandPO |
| 👁️ Visual understanding | MOSS-VL · MOSS-Video-Preview |
| 🎬 Multimodal generation | MOVA · OmniVAE |
| 🌐 Multimodal language models | AnyGPT |
| 🗣️ Speech and audio generation | MOSS-TTS · MOSS-TTS-Nano · MOSS-TTSD · MOSS-Speech · MOSS-Audio-Tokenizer |
| 🎧 Speech, audio, and music understanding | MOSS-Transcribe-Diarize · MOSS-Audio · MOSS-Music |
| 🤖 Embodied AI and robotics | RoboOmni · FRoM-W1 · OpenETA |
| 🔍 Interpretability | Llamascopium (formerly Language-Model-SAEs) · Lorsa |
| 📊 Benchmarks and evaluation | SWE-bench-Science · ContextWeave · AgentHPOBench · FutureOmni · VLABench |
| ⚡ Efficient training and long context | CoLLiE · LongLLaDA · Sparse-dLLM · rope_pp · LongSafety |
| 📚 Surveys and resources | Awesome-WAM · Thus-Spake-Long-Context-LLM |
- SWE-bench-Science: Tests whether coding agents can resolve real engineering issues in scientific software.
- ContextWeave: Evaluates memory systems for coding agents through long-horizon worklog tasks.
- AgentHPOBench: Measures how well LLM agents improve machine learning experiments through sequential hyperparameter changes.
- OmniVAE: Aligns audio and video in a shared latent space for joint reconstruction and generation.
- OpenETA: Connects perception, action, verification, and learning in a continuous physical-world loop.
- MOSS-Transcribe-Diarize: Produces timestamped, speaker-aware transcripts and acoustic event annotations for long recordings.
- MOSS-TTS-Nano: Runs multilingual voice cloning and real-time speech generation on a CPU with 100 million parameters.
- MOSS-VL: Provides open-weight 11B models for long-form and real-time video understanding.
- MOSS-TTS: Covers long-form speech, dialogue synthesis, voice design, sound effects, and streaming TTS.
- MOVA: Generates synchronized video and audio within a single model.
See the pinned repositories for quick access, or browse all 60+ repositories.
For PhD and internship openings, research collaborations, or general inquiries, contact openmoss@sii.edu.cn.