Chinese Folk
Regional styles across provinces and ethnic minorities, with curated scores and descriptions.
Multilingual Folk Music × Multimodal AI
「诗言志,歌永言,声依永,律和声。」 — 《尚书·舜典》
UniVerse(同谣)names a shared verse: musics of many cultures sounding together, humans and machines creating side by side, and past and present inspiring one another — a multilingual folk-music benchmark spanning Chinese folk, Korean gugak, and 38+ languages.
Interactive Gallery
Top to bottom: a rotatable 3D audio map, the bubble chamber for a chosen constellation, then the song card. Drag the map to orbit; hover a point to see its link to a bubble.
Choose a cluster on the map to reveal its bubbles
Waiting for a bubble…
Dataset
Composition of UniVerse across languages, question types, and audio sources.
Regional styles across provinces and ethnic minorities, with curated scores and descriptions.
Arirang lineages and vernacular songs with structured metadata from gugak sources.
Dozens of languages — from Arabic muwashshah to Japanese sakura songs — under one protocol.
Method
Three complementary strategies for multilingual audio–language models: language-balanced SFT, preference optimization on text and audio, and latent reasoning with REPA.
Facing skewed multilingual counts, we keep the natural mix and apply weighted cross-entropy. Language weights are w̃ℓ = min((N / nℓ)α, wmax), then normalized so high-resource languages (e.g. Spanish → ~0.25) down-weight and low-resource ones (e.g. Korean → ~1.76) up-weight. Both the audio tower and the LLM are trained.
Text DPO prefers a chosen response y+ over a rejected y− for fixed audio. Audio DPO freezes the LLM and contrasts positive audio a+ with negative audio a− for a shared response y, so the audio tower must attend to the correct acoustic evidence.
Latent tokens z1…z6 sit between <latent_start> and <latent_end>.
REPA aligns hidden states with teacher vectors from cached audio features (cosine loss).
Phase a fills slots in one forward pass; Phase b refines them recurrently.
Total: Llatent = LLM + λ LREPA.
Main Results
Accuracy (%) on the UniVerse final benchmark (5,042 items), broken down by language group and question type.
| Model | Overall | zh | ko | L-oth. | FE | CR | PR | CA | ED |
|---|
Notes. zh / ko / L-oth.: Chinese, Korean, and other languages. FE: Feature Extraction; CR: Causal Reasoning; PR: Music Pattern Recognition; CA: Comparison Analysis; ED: Error Detection. Qwen2.5 tuned = LR Phase-b; Qwen3 tuned = SFT with thinking.
Per-language scatter: median splits on training utterances and accuracy gain (same quadrants as the language tables below).
Each cell: original | best-tuned. Violet names mark provinces with ≥1 minority-language item.
| Region | Qwen2.5 | Qwen3 | Region | Qwen2.5 | Qwen3 |
|---|
Left: province. Right: melodic mode (tori). Each cell: original | best-tuned.
| Region | Qwen2.5 | Qwen3 | Mode | Qwen2.5 | Qwen3 |
|---|
Cell colors follow train-count vs gain quadrants (green / blue / grey / red).
| Language | Qwen2.5 | Qwen3 | Language | Qwen2.5 | Qwen3 |
|---|