Skip to content
@OpenMOSS

OpenMOSS (SII)

OpenMOSS Team is a research group under the Shanghai Innovation Institution (SII), working in close collaboration with Fudan University and MOSI Intelligence.
OpenMOSS

Shanghai Innovation Institute (SII) · Fudan University · MOSI.AI

Open research on foundation models for language, perception, speech, and embodied intelligence.

Website Hugging Face GitHub Email


👋 About us

OpenMOSS is led by Prof. Xipeng Qiu at the Shanghai Innovation Institute (SII), in collaboration with Fudan University and MOSI.AI. We release models, datasets, benchmarks, and research tools for language, multimodal perception, speech, and embodied intelligence.

🔬 Research directions

Direction Flagship repositories
🧠 Language models MOSS · DiRL · BandPO
👁️ Visual understanding MOSS-VL · MOSS-Video-Preview
🎬 Multimodal generation MOVA · OmniVAE
🌐 Multimodal language models AnyGPT
🗣️ Speech and audio generation MOSS-TTS · MOSS-TTS-Nano · MOSS-TTSD · MOSS-Speech · MOSS-Audio-Tokenizer
🎧 Speech, audio, and music understanding MOSS-Transcribe-Diarize · MOSS-Audio · MOSS-Music
🤖 Embodied AI and robotics RoboOmni · FRoM-W1 · OpenETA
🔍 Interpretability Llamascopium (formerly Language-Model-SAEs) · Lorsa
📊 Benchmarks and evaluation SWE-bench-Science · ContextWeave · AgentHPOBench · FutureOmni · VLABench
Efficient training and long context CoLLiE · LongLLaDA · Sparse-dLLM · rope_pp · LongSafety
📚 Surveys and resources Awesome-WAM · Thus-Spake-Long-Context-LLM

✨ Recent highlights

  • SWE-bench-Science: Tests whether coding agents can resolve real engineering issues in scientific software.
  • ContextWeave: Evaluates memory systems for coding agents through long-horizon worklog tasks.
  • AgentHPOBench: Measures how well LLM agents improve machine learning experiments through sequential hyperparameter changes.
  • OmniVAE: Aligns audio and video in a shared latent space for joint reconstruction and generation.
  • OpenETA: Connects perception, action, verification, and learning in a continuous physical-world loop.
  • MOSS-Transcribe-Diarize: Produces timestamped, speaker-aware transcripts and acoustic event annotations for long recordings.
  • MOSS-TTS-Nano: Runs multilingual voice cloning and real-time speech generation on a CPU with 100 million parameters.
  • MOSS-VL: Provides open-weight 11B models for long-form and real-time video understanding.
  • MOSS-TTS: Covers long-form speech, dialogue synthesis, voice design, sound effects, and streaming TTS.
  • MOVA: Generates synchronized video and audio within a single model.

See the pinned repositories for quick access, or browse all 60+ repositories.

🤝 Join us

For PhD and internship openings, research collaborations, or general inquiries, contact openmoss@sii.edu.cn.

Pinned Loading

  1. MOSS MOSS Public

    An open-source, tool-augmented conversational language model from Fudan University

    Python 12.2k 1.1k

  2. MOSS-VL MOSS-VL Public

    An open-weight 11B model series for long-form and real-time video understanding

    Python 494 17

  3. MOVA MOVA Public

    A foundation model that generates synchronized video and audio in a single model

    Python 1.1k 91

  4. MOSS-TTS-Nano MOSS-TTS-Nano Public

    A 100M-parameter multilingual TTS model for real-time CPU inference, voice cloning, and 48 kHz stereo generation

    Python 4.3k 546

  5. Llamascopium Llamascopium Public

    A framework for training, analyzing, and visualizing sparse autoencoders and related interpretability methods

    Python 227 29

  6. MOSS-Transcribe-Diarize MOSS-Transcribe-Diarize Public

    A 0.9B model for long-form transcription in 50+ languages with speaker diarization, timestamps, and acoustic event awareness

    Python 1.8k 106

Repositories

Showing 10 of 63 repositories

Top languages

Loading…

Most used topics

Loading…