Data Marketplace

Flitto Data Marketplace offers license-verified AI training and evaluation datasets in text, speech, image, and video across 100+ languages and 23+ domains. Try a free sample, then contact us to purchase.

Sell your data on Flitto

Simply register your data and check its sale eligibility.

A total of 326 datasets
  • Fine-tuning DataText

    Korean Telecommunications Domain LLM Question Collection Text Dataset

    A Korean telecommunications domain question collection dataset built for LLM training, with domain-specific questions directly collected and curated by professional annotators.

    DomainIT and Tech
    LanguageKorean
  • Fine-tuning DataText

    Korean QA Text Dataset

    A Korean QA text dataset built from Korean news articles, with related FAQs generated and curated by professional annotators.

    DomainHumanities and Social
    LanguageKorean
  • Pre-training DataText

    EU Privacy Law Legal Document Structuring Text Dataset

    A legal document text dataset built by OCR-processing, hierarchically structuring, and translating EU privacy-related legal documents such as the GDPR and AI Act into Korean and English.

    DomainLaw and Public
    LanguageEnglish|Korean
  • Pre-training DataText

    Low-Resource Language Parallel Translation Corpus Text Dataset

    A low-resource language parallel corpus dataset built by translating Korean written and spoken-style source texts into eight low-resource languages, including Vietnamese, Indonesian, Thai, and Hindi.

    DomainHumanities and Social
    LanguageKorean|English|Hindi|Indonesian|Khmer|Russian|Thai|Uzbek|Vietnamese|Tagalog
  • Frontier DataText

    Gulf Arabic Doctoral Exam Questions and Solutions

    A high-difficulty text dataset developed to train expert-level reasoning in LLMs based on doctoral examination questions and solutions in Gulf Arabic.

    DomainScience and Engineering
    LanguageArabic
  • Frontier DataText

    Egyptian Arabic Doctoral Exam Questions and Solutions

    A high-difficulty text dataset developed to train expert-level reasoning in LLMs based on doctoral examination questions and solutions in Egyptian Arabic.

    DomainScience and Engineering
    LanguageArabic (Egypt)
  • Frontier DataText

    Bengali Doctoral Exam Questions and Solutions

    A high-difficulty text dataset developed to train expert-level reasoning in LLMs based on doctoral examination questions and solutions in Bengali.

    DomainScience and Engineering
    LanguageBengali
  • Frontier DataText

    Hindi Doctoral Exam Questions and Solutions

    A high-difficulty text dataset developed to train expert-level reasoning in LLMs based on doctoral examination questions and solutions in Hindi.

    DomainScience and Engineering
    LanguageHindi
  • Frontier DataText

    Indonesian Doctoral Exam Questions and Solutions

    A high-difficulty text dataset developed to train expert-level reasoning in LLMs based on doctoral examination questions and solutions in Indonesian.

    DomainScience and Engineering
    LanguageIndonesian
  • Frontier DataText

    Japanese Doctoral Exam Questions and Solutions

    A high-difficulty text dataset developed to train expert-level reasoning in LLMs based on doctoral examination questions and solutions in Japanese.

    DomainScience and Engineering
    LanguageJapanese