Multilingual Folk Music × Multimodal AI

One shared verse.
Across cultures, minds, and ages.

「诗言志,歌永言,声依永,律和声。」 — 《尚书·舜典》

UniVerse(同谣)names a shared verse: musics of many cultures sounding together, humans and machines creating side by side, and past and present inspiring one another — a multilingual folk-music benchmark spanning Chinese folk, Korean gugak, and 38+ languages.

  • 1.4MMetadata tokens
  • 38+Languages
  • 5,042QA
  • 1,080Audio min

Interactive Gallery

Sound Map & Bubble Museum

Top to bottom: a rotatable 3D audio map, the bubble chamber for a chosen constellation, then the song card. Drag the map to orbit; hover a point to see its link to a bubble.

1 · Global 3D audio map

Drag to rotate · scroll to zoom · click a point or constellation. Hover a point to reveal its bubble link.

2 · Bubble chamber

Choose a cluster on the map to reveal its bubbles

3 · Exhibit

Pop a bubble to read the song card here.

Waiting for a bubble…

Dataset

Benchmark at a Glance

Composition of UniVerse across languages, question types, and audio sources.

Chinese Folk

Regional styles across provinces and ethnic minorities, with curated scores and descriptions.

Korean Gugak

Arirang lineages and vernacular songs with structured metadata from gugak sources.

World Folk

Dozens of languages — from Arabic muwashshah to Japanese sakura songs — under one protocol.

Benchmark composition pie charts: language, question type, and audio source
Figure — Benchmark composition (language / question type / audio source).

Method

Training Strategies

Three complementary strategies for multilingual audio–language models: language-balanced SFT, preference optimization on text and audio, and latent reasoning with REPA.

Training strategies overview: Lang Loss SFT, Text and Audio DPO, and Latent Reasoning with REPA
Figure — Training strategies from the paper (Lang Loss SFT, Text & Audio DPO, Latent Reasoning + REPA).

Lang. Loss SFT

Facing skewed multilingual counts, we keep the natural mix and apply weighted cross-entropy. Language weights are = min((N / n)α, wmax), then normalized so high-resource languages (e.g. Spanish → ~0.25) down-weight and low-resource ones (e.g. Korean → ~1.76) up-weight. Both the audio tower and the LLM are trained.

Text & Audio DPO

Text DPO prefers a chosen response y+ over a rejected y for fixed audio. Audio DPO freezes the LLM and contrasts positive audio a+ with negative audio a for a shared response y, so the audio tower must attend to the correct acoustic evidence.

Latent Reasoning + REPA

Latent tokens z1…z6 sit between <latent_start> and <latent_end>. REPA aligns hidden states with teacher vectors from cached audio features (cosine loss). Phase a fills slots in one forward pass; Phase b refines them recurrently. Total: Llatent = LLM + λ LREPA.

Main Results

Overall Results

Accuracy (%) on the UniVerse final benchmark (5,042 items), broken down by language group and question type.

Model Overall zh ko L-oth. FE CR PR CA ED

Notes. zh / ko / L-oth.: Chinese, Korean, and other languages. FE: Feature Extraction; CR: Causal Reasoning; PR: Music Pattern Recognition; CA: Comparison Analysis; ED: Error Detection. Qwen2.5 tuned = LR Phase-b; Qwen3 tuned = SFT with thinking.

Training Count vs Tuning Gain

Per-language scatter: median splits on training utterances and accuracy gain (same quadrants as the language tables below).

Scatter plot of training utterance count versus tuning gain by language for Qwen2.5 and Qwen3
Figure — Samples of training data vs tuning gain by language (Qwen2.5 LR Phase-b; Qwen3 SFT).

Chinese Regions

Each cell: original | best-tuned. Violet names mark provinces with ≥1 minority-language item.

RegionQwen2.5Qwen3 RegionQwen2.5Qwen3

Korean Provinces & Tori

Left: province. Right: melodic mode (tori). Each cell: original | best-tuned.

RegionQwen2.5Qwen3 ModeQwen2.5Qwen3

Foreign Languages (Low-Resource Subset)

Cell colors follow train-count vs gain quadrants (green / blue / grey / red).

LanguageQwen2.5Qwen3 LanguageQwen2.5Qwen3