General BFCL 4v Benchmark Dataset
A Korean-based BFCL 4v benchmark text dataset built for evaluating function-calling capabilities of AI models.
Use verified, licensed data with confidence. You can download right away or check the data through inquiry.
A Korean-based BFCL 4v benchmark text dataset built for evaluating function-calling capabilities of AI models.
A Korean-based Tau benchmark text dataset built for evaluating AI model performance in benchmark tasks.
A Korean-based Chat Arena preference evaluation single-turn dataset built by collecting human preferences for AI model responses.
A Korean-based Chat Arena preference evaluation multi turn dataset built by collecting human preferences for AI model responses.
A Korean-based long context benchmark text dataset built to evaluate AI models' ability to process and reason over extended contexts.
A high-difficulty Korean reasoning benchmark dataset built according to global HLE standards, including complex expert-level reasoning problems.
A high-difficulty Korean HLE reasoning and benchmark dataset built with Korea-specific expert-level reasoning problems.
A benchmark text dataset developed for Arena evaluation based on Gulf Arabic.
A benchmark text dataset developed for Arena evaluation based on Egyptian Arabic.
A benchmark text dataset developed for Arena evaluation based on Bengali.