Skip to content
@Toloka

Toloka

Data labeling platform for ML

Pinned Loading

  1. tolokaforge tolokaforge Public

    Universal LLM benchmarking harness for tool use, browser, mobile, coding, and long-horizon evals

    Python 17 8

  2. tendem-evaluation tendem-evaluation Public

    Tendem hybrid AI+Human system benchmarking

    Python 3

  3. beemo beemo Public

    Benchmark for fine-grained machine-generated text detection. 6.5k texts written by humans, generated by ten open-source instruction-finetuned LLMs and edited by expert annotators.

    11 2

  4. u-math u-math Public

    Official evaluation code for the U-MATH and μ-MATH benchmarks. These datasets are designed to test the mathematical reasoning and meta-evaluation capabilities of LLMs on university-level problems.

    Python 11 3

  5. crowd-kit crowd-kit Public

    Control the quality of your labeled data with the Python tools you already know.

    Python 254 21

Repositories

Showing 10 of 33 repositories
  • tolokaforge Public

    Universal LLM benchmarking harness for tool use, browser, mobile, coding, and long-horizon evals

    Toloka/tolokaforge's past year of commit activity
    Python 17 8 456 12 Updated Sep 11, 2026
  • tendem-mcp Public

    Home for Human In The Loop Tendem plugin for your agent.

    Toloka/tendem-mcp's past year of commit activity
    Python 102 MIT 6 2 1 Updated Sep 8, 2026
  • Toloka/n8n-nodes-tendem's past year of commit activity
    TypeScript 0 MIT 0 1 0 Updated Aug 27, 2026
  • Toloka/template-builder's past year of commit activity
    TypeScript 4 Apache-2.0 1 1 7 Updated May 28, 2026
  • dbxio Public

    High-level Databricks client

    Toloka/dbxio's past year of commit activity
    Python 13 0 1 5 Updated May 19, 2026
  • CrowdSpeech Public

    Benchmark Dataset for Crowdsourced Audio Transcription

    Toloka/CrowdSpeech's past year of commit activity
    Python 11 3 1 6 Updated Apr 18, 2026
  • dbt-af Public

    Distributed run of dbt models using Airflow

    Toloka/dbt-af's past year of commit activity
    Python 168 14 1 2 Updated Apr 14, 2026
  • pg-queue-playground Public

    Playground for transactional queues in PostgreSQL

    Toloka/pg-queue-playground's past year of commit activity
    Java 6 0 2 1 Updated Apr 12, 2026
  • u-math Public

    Official evaluation code for the U-MATH and μ-MATH benchmarks. These datasets are designed to test the mathematical reasoning and meta-evaluation capabilities of LLMs on university-level problems.

    Toloka/u-math's past year of commit activity
    Python 11 MIT 3 1 0 Updated Jan 30, 2026
  • tendem-evaluation Public

    Tendem hybrid AI+Human system benchmarking

    Toloka/tendem-evaluation's past year of commit activity
    Python 3 0 0 0 Updated Jan 27, 2026

Top languages

Loading…