SapinSapin AI — Open Foundations for Philippine AI
SapinSapin AI builds open speech and language foundations so Philippine AI can be made with — and for — the people who speak it.
What We Build
Open datasets, speech recognition models, text-to-speech systems, and language models for 10+ Philippine languages: Filipino, Cebuano, Ilocano, Hiligaynon, Bikol, Pangasinan, Kapampangan, Tausug, Waray, and English.
Datasets
- Philippine Language Dataset (PLD) — 334,268 utterances across 10 languages, 448+ hours of prompted 16 kHz speech
- Filipino Speech Corpus — 305,246 segments, 65+ hours of Filipino/Tagalog speech (MIT license)
- BantayWika — 6.09M verified tokens of Filipino, Cebuano, and Ilocano text
- halohalo — 45,536 rows of cleaned Philippine-language web data
Models
28 public models spanning speech recognition (Whisper fine-tunes), text-to-speech (SpeechT5 fine-tunes), voice conversion, and text generation (Llama 3.1 and GPT-oss-20b continued pretraining).
Live Demo
Try speech recognition, text-to-speech, and voice conversion for Philippine languages at the halohalo Space on Hugging Face.
FAQ
- Which license applies?
- Licenses are dataset-specific. MIT and the UP-DSP research license appear among the current datasets. Always read the dataset card before use.
- Can I use the data commercially?
- It depends on the individual dataset license and any access conditions. The linked Hugging Face dataset card is the source of record.
- How do I contribute?
- Open an issue or discussion on a SapinSapin repository on GitHub.
- What is the model roadmap?
- The catalog includes language, speech recognition, text-to-speech, and audio-to-audio models. Follow the Hugging Face activity for updates.
Links
© 2026 SapinSapin AI · Designed for open research