Data Marketplace

Use verified, licensed data with confidence. You can download right away or check the data through inquiry.

Sell your data on Flitto

Simply register your data and check its sale eligibility.

A total of 326 datasets
  • Pre-training DataAudio

    English Speech Dataset by Korean Speakers

    A single-turn English speech dataset recorded by Korean speakers.

  • Pre-training DataAudio

    Indonesian Speech Dataset by Korean Speakers

    A single-turn Indonesian speech dataset recorded by Korean speakers.

  • Pre-training DataAudio

    Vietnamese Speech Dataset by Korean Speakers

    A single-turn Vietnamese speech dataset recorded by Korean speakers.

  • Pre-training DataAudio

    English Finance Domain Speech Dataset

    A single-turn English speech dataset specialized for the finance domain.

  • Pre-training DataAudio

    English Single-Utterance Speech Dataset for ICT & Software Domains

    A single-turn English speech dataset featuring monologues by IT and software professionals about their work experiences.

  • Pre-training DataAudio

    English Single-Utterance Speech Dataset for Finance & Economics Domains

    A single-turn English speech dataset featuring monologues about workplace experiences in finance.

  • Fine-tuning DataText

    Python, JavaScript, and Java Code Generation and Refactoring Dataset

    A code generation and refactoring dataset built from source code snippets, with software and IT specialists creating improvement instructions and revised code outputs.

  • Fine-tuning DataText

    Korean Code Instruction and Commenting Text Dataset

    A Korean code completion dataset built from source code snippets, Korean declarative and imperative instructions, and block-level and line-level comments written by software and IT specialists.

  • Alignment DataText

    English Safety Multi-Turn Chat Text Dataset

    An English safety multi-turn chat text dataset built by professional annotators for safety alignment training based on defined safety criteria.

  • Pre-training DataText

    Arabic General-Domain Multi-Turn Chat Text Dataset

    An Arabic general-domain multi-turn chat text dataset built by professional native-language annotators for daily life and general conversation domains.