Why This Page Exists
Anyone who sets out to build a Filipino speech recognition system will, sooner or later, run into the same question: where is the data?
The answer is more complicated than it should be. Filipino speech datasets exist. Some are substantial. But they are scattered across institutional repositories, government-funded projects, commercial vendors, and research challenge archives, each with its own licensing terms, label conventions, and undocumented quirks. A researcher in Manila and a product team in San Francisco will waste the same weeks discovering what is available, negotiating access, and learning the hard way that a dataset’s documentation does not always match its contents.
This page is an attempt to save that time. It catalogs the Filipino and Tagalog speech datasets I am aware of, with honest notes on what each one actually contains, how usable the labels are, and what to watch out for. It is not exhaustive. Researchers and organizations with datasets not listed here are welcome to reach out.
The datasets are grouped into three categories: labeled datasets suitable for training or evaluation, unlabeled or weakly labeled datasets useful for self-supervised pretraining or semi-supervised approaches, and commercial datasets available through vendors.
Labeled Filipino Speech Datasets
These datasets have transcription labels and are, in principle, usable for supervised ASR training or evaluation. The hours listed are rough estimates of transcribed speech, excluding silence. Some datasets lack established train/dev/test splits; where no split exists, the hours are listed under training.
| Dataset | Domain | Style | Train (h) | Dev (h) | Test (h) | Transcription Quality | License |
|---|---|---|---|---|---|---|---|
| Filipino Speech Corpus (FSC) vols. 1–4 | Mixed prompts | Read | ~52.5 | — | — | Normalized, uneven | ELRA User Agreement |
| PLD (Filipino subset, formerly ISIP Corpus) | News, medical, education, tourism, spontaneous | Read, spontaneous | ~33.5 (Filipino subset) | — | — | Normalized, incomplete | Institutional (Mozilla Data Collective) |
| CMU-KIT Filipino Corpus | MT basic/travel/medical, news | Read, semi-spontaneous | ~28 | — | — | Normalized | Institutional |
| IARPA BABEL Tagalog | Conversational, scripted calls | Spontaneous, read | ~115 | ~13 | ~85 (restricted) | Punctuated, cased | LDC User Agreement |
| IARPA MATERIAL (excl. BABEL) | Broadcast news | Spontaneous | ~7 | ~7 | ~3.5 | Punctuated, cased | NIST Agreement |
| FLEURS fil_ph | Wikipedia (translated) | Read | ~8 | ~2 | ~5 | Normalized; punctuated, cased | CC-BY-4.0 |
| MagicHub ASR-SFDuSC | Daily-use sentences | Read | ~4.6 | — | — | Transcribed | Open (MagicHub) |
Notes on Each Dataset
Filipino Speech Corpus (FSC), volumes 1 through 4. This was the earliest organized attempt at a Filipino speech corpus, developed at the University of the Philippines. Despite the name, the prompts are almost entirely in Tagalog. The corpus includes syllable sequences intended for TTS development, not ASR. Recordings were made in long continuous sessions, making precise alignment dependent on timestamps that were not always reliable. The designers invested heavily in documentation and file-naming conventions (with placeholder variables for future classification schemes) but less in label quality. Phonetic coverage is narrow: many participants read identical sentences, which makes it difficult to construct non-overlapping train/dev/test splits. The metadata and labels were poorly managed across volumes, and the net result is data that is partially usable but requires significant cleaning.
Philippine Languages Database (PLD), formerly the ISIP Corpus. The PLD was presented at SIGUL @ LREC-COLING 2024 (Guevara, Cajote, Bayona, and Lucas) as a multilingual corpus covering ten Philippine languages with over 454 hours total. The Filipino subset originates from data collected under the government-funded Interdisciplinary Signal Processing for Pinoys (ISIP) program around 2010 at the University of the Philippines. The corpus shares several characteristics with the FSC: limited prompt variety, no established train/dev/test splits, and incomplete transcriptions. The final three questions in each recording session were answered spontaneously but never transcribed, leaving that portion of the speech unlabeled (and excluded from the hours listed above). The broader PLD covers additional Philippine languages (Cebuano, Kapampangan, Hiligaynon, Ilokano, Bikolano, Waray, and Tausug) across domains including news, medical, education, tourism, and spontaneous speech, and is downloadable through the Mozilla Data Collective. A quality assessment of the full PLD is a planned future post on this site.
CMU-KIT Filipino Corpus. Created in collaboration between Carnegie Mellon University and the Karlsruhe Institute of Technology around 2008 to 2009, this corpus used machine-translated English prompts from standard NLP resources covering basic expressions, travel, and medical fields, supplemented by news articles. The machine translations often produced archaic or transliterated Tagalog, and recorders were not permitted to modify the prompts. This means the read speech reflects unnatural sentence constructions that no native speaker would produce unprompted. A small spontaneous component exists but is limited in scope.
IARPA BABEL Tagalog. The most substantial labeled Filipino dataset publicly documented, at roughly 213 hours total (115 training, with dev and a restricted test set accessible through the openKWS evaluation framework). Produced by Appen for the BABEL program in 2012. Despite being labeled “Tagalog,” the content is functionally Filipino: conversations include code-switching, Manila-register vocabulary, and the full range of informal spoken Filipino. The critical caveat is channel quality. Recordings were made over mobile wireless connections, resulting in intermittent dropouts, muffled narrowband audio, and choppy segments. For ASR training this is a feature as much as a limitation (models trained on this data learn to handle degraded channels), but for evaluation it means WER scores on BABEL are not comparable to scores on clean read-speech datasets. The transcription labels are punctuated and cased, with disfluency tags (such as <hes>) and overlap annotations. This labeling standard, inherited from the Hub-4 and Fisher conventions, is richer than what most modern datasets provide.
IARPA MATERIAL. A follow-on program to BABEL, launched between 2016 and 2017, also including Filipino. Some of the audio derives from BABEL itself; the remainder comes from broadcast news sources, likely YouTube. The non-BABEL portion is smaller but covers a different domain (produced media rather than telephone conversations). Labels follow the same punctuated, cased convention as BABEL.
FLEURS fil_ph. The dataset most commonly used to benchmark Filipino ASR, and the one with the most serious structural problems. FLEURS is built on FLoRes, a machine translation benchmark: 2,009 English Wikipedia sentences translated by professional translators into 102 languages, then recorded by native speakers. Every sentence in the Filipino subset is a translation. Not one originated in Filipino. This introduces translationese bias: sentence structures that follow English SVO patterns rather than Filipino’s natural VSO and topic-comment constructions, vocabulary choices shaped by the source text rather than by how a Filipino speaker would express the same idea, and loanword spellings determined by the translator rather than by convention. A 2025 Google audit of FLEURS, Common Voice, and VoxPopuli found that quality flaws create an “illusion of success” in low-resource languages, with mislabeled language varieties inflating substitution rates by as much as 25 percentage points in structurally similar cases. A systematic audit of 100 FLEURS fil_ph utterances (documented separately in Problems with FLEURS as a Filipino Benchmark) identified nine categories of reference quality issues, from code-switching orthography to speaker misreads that the reference text does not acknowledge. Despite these problems, FLEURS remains the only freely available Filipino speech dataset with an established train/dev/test split, which is why it continues to be used and why its limitations matter.
MagicHub ASR-SFDuSC. A small (4.58 hours, 10 speakers) open-source dataset of scripted daily-use Filipino sentences from Magic Data Technology. Limited in scale but freely accessible and cleanly transcribed, making it potentially useful for quick prototyping or as supplementary fine-tuning data.
Unlabeled and Weakly Labeled Datasets
These datasets contain Filipino speech without reliable transcription labels, or with labels derived from sources other than manual annotation (such as YouTube auto-captions or subtitle files). They are relevant for self-supervised pretraining, semi-supervised training, and weak-label approaches.
| Dataset | Domain | Style | Hours (est.) | Label Status | License |
|---|---|---|---|---|---|
| FSC vol. 5 | Mixed topics | Spontaneous (aware) | ~6.5 | Unlabeled | ELRA User Agreement |
| ICE Philippines | Mixed topics | Oratory, spontaneous | ~29.5 | RTF transcripts exist, poorly formatted | Institutional |
| VoxLingua107 (Filipino subset) | YouTube | Mixed | ~93 | Language ID labels only | CC-BY-4.0 |
| YouTube-8M (Filipino subset) | TV programs | Narrated, acted, spontaneous | ~39 | Video-level tags only | CC-BY-4.0 |
Notes
FSC volume 5 is the spontaneous speech component of the Filipino Speech Corpus. Speakers were aware they were being recorded but spoke freely on chosen topics. The absence of transcriptions makes it useful only for unsupervised methods or as evaluation material if manually transcribed.
ICE Philippines (International Corpus of English, Philippine component) is a peculiar case. Transcriptions do exist in rich-text format, but parsing them into aligned label files requires non-trivial effort that, to my knowledge, nobody has published tooling for. The audio is distributed in lossy MP3, which introduces compression artifacts. The content itself is valuable: oratory, interviews, and spontaneous conversation in Philippine English and Filipino.
VoxLingua107 and YouTube-8M are large-scale datasets originally built for language identification and video classification, respectively. Their Filipino subsets contain substantial hours of speech but no utterance-level transcriptions. For weak-label ASR training, these are source material: the audio is there, and labels must be generated through existing models or manual effort.
Beyond these publicly documented corpora, substantial Filipino speech data can be collected from YouTube through channel-specific crawling. The practical challenge is not collection but labeling. Subtitle files, where they exist, are often auto-generated or non-verbatim (reflecting editorial choices rather than literal transcription). Recent work on robust training methods, including the Weakly Supervised Transducer (WST), has shown that transducer-based ASR can maintain performance even with transcript error rates as high as 70%, which materially changes the calculus for using subtitle-derived training data.
Commercial and Restricted Datasets
These datasets are documented publicly but require commercial licensing, institutional agreements, or challenge registration to access.
| Dataset | Domain | Style | Hours | Speakers | Access |
|---|---|---|---|---|---|
| MagicHub ASR-SFTagaSC | Daily-use sentences | Scripted | 417 | 444 | Commercial (contact vendor) |
| MagicHub ASR-BigFTagaCSC | Topic-based conversation | Conversational | 1,285 | 514 | Commercial (contact vendor) |
| Defined.ai Tagalog Spontaneous Dialogue | Banking, insurance, retail, telecom | Spontaneous dialogue | 224 | Unknown | Commercial |
| Speech-data.ai Filipino Dataset | Mixed | Mixed | 75 | 639 files | Commercial (sample on HF) |
| MLC-SLM 2026 (Tagalog portion) | Conversational | Two-speaker dialogue | TBD | TBD | Challenge registration (free) |
Notes
MagicHub offers the largest documented Filipino speech datasets by raw hours. The 1,285-hour conversational corpus alone dwarfs every other entry on this page. I have not used these datasets and cannot comment on label quality, prompt design, or recording conditions. The scripted corpus uses daily-use sentences, which raises the same coverage questions that affected the FSC: how phonetically diverse are the prompts, and how much speaker variation exists across the 444 contributors? These are questions a buyer should ask before licensing.
Defined.ai’s Tagalog Spontaneous Dialogue dataset is notable for its domain coverage (banking, insurance, retail, telecommunications) and its recording at 8kHz, which signals that the data was collected over telephone channels or intentionally downsampled to match telephony conditions. For contact center ASR applications this is directly relevant.
MLC-SLM 2026 is the second edition of the Multilingual Conversational Speech Language Model Challenge, which added Tagalog in 2026 alongside Urdu, Turkish, and regional variants of French and Spanish. The dataset is free to registered participants and provides two-speaker conversational speech at 16kHz with oracle segmentation and speaker labels. The first edition (2025, Interspeech satellite event) covered 11 languages and attracted 78 teams. For Filipino ASR researchers, this challenge represents a rare opportunity to work with professionally segmented conversational data under a free license, and to benchmark against international participants.
What Is Missing
The landscape described above has several conspicuous gaps.
No Common Voice Tagalog. Despite Common Voice’s coverage of over 100 languages, there is no Tagalog or Filipino subset. The Pashto Common Voice effort, which built a corpus from zero community infrastructure to a usable dataset across several release cycles, offers a template for what a Filipino Common Voice initiative might look like. The absence is not technical; it is organizational.
No open, large-scale, natively Filipino evaluation set. FLEURS is translated. BABEL’s test set is restricted. MATERIAL is small. There is no freely available, natively authored Filipino speech evaluation dataset with established splits, clean labels, and sufficient size to serve as a credible benchmark. This gap is the central problem of Filipino ASR evaluation today. Building or curating such a dataset is arguably the single most valuable contribution the community could make.
No standardized normalization or scoring conventions. Even where labeled data exists, there is no agreed-upon text normalization standard for Filipino ASR evaluation. Orthographic variants (kanyang vs. kaniyang, ngunit vs. nguni’t, bawat vs. bawa’t), code-switched loanword handling, numeral rendering, and disfluency treatment are all handled ad hoc by individual researchers. The scoring methodology essay (How Filipino ASR Should Be Scored) addresses this problem in detail.
Limited spontaneous and conversational data under open licenses. The datasets with the most natural speech (BABEL, Defined.ai, MagicHub conversational) are either paywalled or access-restricted. The freely available datasets (FLEURS, FSC, VoxLingua107) are either read speech, translated speech, or unlabeled. This imbalance means that open Filipino ASR research is biased toward read-speech performance, which does not predict how well a system will handle real conversations.
A Note on Tagalog vs. Filipino
Several datasets in this list use “Tagalog” and “Filipino” interchangeably. The linguistic relationship between the two is complex and politically charged, but for practical ASR purposes the distinction matters less than the register and domain of the speech. BABEL’s “Tagalog” data is informal Manila Filipino with heavy code-switching. FLEURS’s “Filipino” data is translated encyclopedic prose read aloud. The language code (tl for Tagalog, fil for Filipino) tells you less about what the model will encounter than the dataset’s documentation, and sometimes less than the documentation tells you.
This page is a living reference. If you maintain or have access to a Filipino speech dataset not listed here, or if any of the information above is inaccurate, please get in touch. Last updated: August 2026.