Data Marketplace

Flitto Data Marketplace offers license-verified AI training and evaluation datasets in text, speech, image, and video across 100+ languages and 23+ domains. Try a free sample, then contact us to purchase.

Sell your data on Flitto

Simply register your data and check its sale eligibility.

A total of 326 datasets
  • Frontier DataText

    Gulf Arabic Expert CoT Dataset

    A text dataset for LLM Chain-of-Thought training developed on the basis of expert utterances in Gulf Arabic.

    DomainAd and Marketing|Sports, Arts and Culture|Games|Humanities and Social|IT and Tech|Law and Public|Medical|Management, Economic and Finance
    LanguageArabic
  • Frontier DataText

    Egyptian Arabic Expert CoT Dataset

    A text dataset for LLM Chain-of-Thought training developed on the basis of expert utterances in Egyptian Arabic.

    DomainAd and Marketing|Sports, Arts and Culture|Games|Humanities and Social|IT and Tech|Law and Public|Medical|Management, Economic and Finance
    LanguageArabic (Egypt)
  • Frontier DataText

    Bengali Expert CoT Dataset

    A text dataset for LLM Chain-of-Thought training developed on the basis of expert utterances in Bengali.

    DomainAd and Marketing|Sports, Arts and Culture|Games|Humanities and Social|IT and Tech|Law and Public|Medical|Management, Economic and Finance
    LanguageBengali
  • Frontier DataText

    Hindi Expert CoT Dataset

    A text dataset for LLM Chain-of-Thought training developed on the basis of expert utterances in Hindi.

    DomainAd and Marketing|Sports, Arts and Culture|Games|Humanities and Social|IT and Tech|Law and Public|Medical|Management, Economic and Finance
    LanguageHindi
  • Frontier DataText

    Indonesian Expert CoT Dataset

    A text dataset for LLM Chain-of-Thought training developed on the basis of expert utterances in Indonesian.

    DomainAd and Marketing|Sports, Arts and Culture|Games|Humanities and Social|IT and Tech|Law and Public|Medical|Management, Economic and Finance
    LanguageIndonesian
  • Frontier DataText

    Japanese Expert CoT Dataset

    A text dataset for LLM Chain-of-Thought training developed on the basis of expert utterances in Japanese.

    DomainAd and Marketing|Sports, Arts and Culture|Games|Humanities and Social|IT and Tech|Law and Public|Medical|Management, Economic and Finance
    LanguageJapanese
  • Frontier DataText

    Thai Expert CoT Dataset

    A text dataset for LLM Chain-of-Thought training developed on the basis of expert utterances in Thai.

    DomainAd and Marketing|Sports, Arts and Culture|Games|Humanities and Social|IT and Tech|Law and Public|Medical|Management, Economic and Finance
    LanguageThai
  • Alignment DataText

    Korean RAG Instruction Text Dataset

    A Korean RAG instruction text dataset built from documents with user queries, query type labels, answers, and supporting evidence spans written in Korean.

    DomainHumanities and Social
    LanguageKorean
  • Pre-training DataAudio

    Japanese Speech Dataset by Korean Speakers

    A single-turn Japanese speech dataset recorded by Korean speakers.

    DomainHumanities and Social|Lifestyle
    LanguageJapanese
  • Pre-training DataAudio

    Chinese Speech Dataset by Korean Speakers

    A single-turn Chinese speech dataset recorded by Korean speakers.

    DomainHumanities and Social|Lifestyle
    LanguageChinese (Simplified)