Data Marketplace

Flitto Data Marketplace offers license-verified AI training and evaluation datasets in text, speech, image, and video across 100+ languages and 23+ domains. Try a free sample, then contact us to purchase.

Sell your data on Flitto

Simply register your data and check its sale eligibility.

A total of 326 datasets
  • Pre-training DataImage

    Arabic Food Menu and Handwritten OCR Image Dataset

    An Arabic OCR image dataset for real-world food menu and printed document images, built with JSON annotations containing bounding boxes and language metadata.

    DomainLifestyle
    LanguageArabic
  • Pre-training DataImage

    Arabic Magazine Cover OCR Image Dataset

    An Arabic OCR image dataset for real-world magazine cover images in the media and entertainment domain, built with JSON annotations containing bounding-box coordinates.

    DomainAd and Marketing
    LanguageArabic
  • Pre-training DataImage

    Arabic Advertising Poster and Flyer OCR Image Dataset

    An Arabic OCR image dataset for real-world posters and flyers in the media and retail domains, built with JSON annotations containing bounding boxes and language metadata.

    DomainAd and Marketing
    LanguageArabic
  • Pre-training DataImage

    Arabic Product Packaging Back-Side OCR Image Dataset

    An Arabic OCR image dataset for real-world back-side product packaging images in the retail and food domains, built with JSON annotations containing bounding-box coordinates.

    DomainLifestyle
    LanguageArabic
  • Pre-training DataImage

    Arabic Product Packaging Front-Side OCR Image Dataset

    An Arabic OCR image dataset for real-world front-side product packaging images in the retail and food domains, built with JSON annotations containing bounding-box coordinates.

    DomainLifestyle
    LanguageArabic
  • Pre-training DataImage

    Arabic Vowelized Text OCR Image Dataset

    An Arabic OCR image dataset for real-world vowelized text images in education and general domains, built with JSON annotations containing bounding boxes and language metadata.

    DomainEducation
    LanguageArabic
  • Pre-training DataImage

    Ukrainian Printed Text OCR Image Dataset

    A Ukrainian printed text OCR image dataset built from real-world images with JSON annotations containing bounding boxes, language metadata, and reading direction information.

    DomainHumanities and Social
    LanguageUkrainian
  • Pre-training DataImage

    Ukrainian Handwritten Text OCR Image Dataset

    A Ukrainian handwritten text OCR image dataset built from real-world images with JSON annotations containing bounding boxes, language metadata, and reading direction information.

    DomainHumanities and Social
    LanguageUkrainian
  • Pre-training DataImage

    Japanese Handwritten Text OCR Image Dataset

    A Japanese handwritten text OCR image dataset built from real-world images with JSON annotations containing bounding boxes, language metadata, and reading direction information.

    DomainHumanities and Social
    LanguageJapanese
  • Pre-training DataImage

    Thai Printed Text OCR Image Dataset

    A Thai printed text OCR image dataset built from real-world images with JSON annotations containing bounding boxes, language metadata, and reading direction information.

    DomainHumanities and Social
    LanguageThai