Data Marketplace

Flitto Data Marketplace offers license-verified AI training and evaluation datasets in text, speech, image, and video across 100+ languages and 23+ domains. Try a free sample, then contact us to purchase.

Sell your data on Flitto

Simply register your data and check its sale eligibility.

A total of 326 datasets
  • Pre-training DataAudio

    Korean Multi-Speaker Multi-Turn Speech Dataset with 3 to 4+ Speakers

    A natural Korean multi-turn spoken dialogue dataset featuring conversations among three to four or more Korean speakers.

    DomainHumanities and Social|Lifestyle
    LanguageKorean
  • Pre-training DataAudio

    English Multi-Speaker Multi-Turn Speech Dataset with 3 to 4+ Speakers

    A natural English multi-turn spoken dialogue dataset featuring conversations among three to four or more native English speakers.

    DomainHumanities and Social|Lifestyle
    LanguageEnglish
  • Pre-training DataAudio

    French Vehicle Control Command Speech Dataset for Multilingual Applications

    A vehicle control command speech dataset designed to enhance in-vehicle voice interfaces.

    DomainIT and Tech
    LanguageEnglish|French (Canada)|Spanish (Latin America)
  • Pre-training DataAudio

    Korean Robot Work Instruction Speech Dataset for Industrial Sites

    A voice command dataset for instructing robots or automated equipment in manufacturing, logistics, and industrial environments.

    DomainIT and Tech
    LanguageKorean
  • Pre-training DataAudio

    Korean Call Center Consultation Speech Dataset with Emotion Labels

    A call center consultation speech dataset based on customer-agent conversations and annotated with emotion labels.

    DomainHumanities and Social
    LanguageKorean
  • Pre-training DataAudio

    English Call Center Consultation Speech Dataset with Emotion Labels

    A call center consultation speech dataset based on customer-agent conversations and annotated with emotion labels.

    DomainHumanities and Social
    LanguageEnglish
  • Pre-training DataAudio

    Code-Switching Multi-Turn Speech Dataset

    A code-switching multi-turn spoken dialogue dataset in which two or more languages are naturally mixed within a single conversation.

    DomainHumanities and Social
    LanguageKorean|English|Japanese|Chinese (Simplified)|Spanish|Arabic
  • Pre-training DataAudio

    Emergency Voice Command Dataset

    A dataset of short voice commands and warning turns used in emergency situations such as disasters, accidents, security incidents, medical emergencies, and defense scenarios.

    DomainLaw and Public
    LanguageKorean
  • Pre-training DataImage

    Arabic Retail Storefront Sign OCR Image Dataset

    An Arabic OCR image dataset for real-world retail storefront signs, built with JSON annotations containing bounding boxes, language metadata, and reading direction information.

    DomainLifestyle
    LanguageArabic
  • Pre-training DataImage

    Arabic Public Transportation Sign and Notice OCR Image Dataset

    An Arabic OCR image dataset for real-world public and transportation signs and notices, built with JSON annotations containing bounding boxes and language metadata.

    DomainLaw and Public
    LanguageArabic