From 547a2a5f92616b17fd6fcfd950828c5a1de66b64 Mon Sep 17 00:00:00 2001 From: Victor Skvortsov Date: Wed, 7 Oct 2026 17:10:22 +0500 Subject: [PATCH 1/2] Support trace-replay datasets such as AIPerf AgentX in presets --- mkdocs/docs/concepts/presets.md | 13 ++++++++++ .../presets/resources/system_prompt.md | 26 ++++++++++++++++--- .../_internal/core/models/configurations.py | 3 ++- .../cli/services/presets/test_prompt.py | 4 +++ 4 files changed, 42 insertions(+), 4 deletions(-) diff --git a/mkdocs/docs/concepts/presets.md b/mkdocs/docs/concepts/presets.md index 5a0ef995da..8a2b12e905 100644 --- a/mkdocs/docs/concepts/presets.md +++ b/mkdocs/docs/concepts/presets.md @@ -216,6 +216,19 @@ The requests every benchmark measures. The dataset provides the requests, so `input_tokens`, `output_tokens`, and `shared_prefix_tokens` can't be set with it, and the preset records the measured means. A gated dataset requires `HF_TOKEN` in `env`. + `dataset` can also name a dataset of recorded multi-turn sessions, such as one of the public agentic trace datasets of [SemiAnalysis's AIPerf fork](https://github.com/SemiAnalysisAI/aiperf). The agent then replays the sessions with their turn order and timing, so that prefix reuse matches real traffic: + + ```yaml + dataset: semianalysis_cc_traces_weka_062126_256k + concurrency: 16 + + prompt: | + Benchmark with AIPerf `--scenario inferencex-agentx-mvp`. Use a 600-second + replay for trials and a 3600-second replay for the final benchmark. + ``` + + A replay runs for a fixed duration rather than a number of requests, so creation takes much longer than with synthetic prompts. Use `prompt` to set the replay options, such as the duration. + ### Baseline By default, the first trial is a baseline: the agent serves the model the way the chosen serving framework recommends, without tuning it for performance. Later trials are optimization attempts. Set `baseline: false` to make every trial an optimization attempt. diff --git a/src/dstack/_internal/cli/services/presets/resources/system_prompt.md b/src/dstack/_internal/cli/services/presets/resources/system_prompt.md index 691df34b33..d367b8f799 100644 --- a/src/dstack/_internal/cli/services/presets/resources/system_prompt.md +++ b/src/dstack/_internal/cli/services/presets/resources/system_prompt.md @@ -342,10 +342,21 @@ example, the dataset options are: | --- | --- | | `vllm bench serve` | `--dataset-name ` when `dataset` is the tool's own dataset name, or `--dataset-name hf --dataset-path ` when it is a Hugging Face dataset ID | | `sglang.benchmark.serving` | `--dataset-name ` when `dataset` is the tool's own dataset name; the tool has no Hugging Face dataset option | +| `aiperf profile` | `--public-dataset ` when `dataset` is one of AIPerf's public datasets; for agentic trace datasets such as `semianalysis_cc_traces_weka_062126_256k`, also `--scenario inferencex-agentx-mvp` (SemiAnalysis's AIPerf fork), which replays multi-turn sessions with their timing | The table is an example and not a full command: the remaining options still come from `concurrency`, option names and defaults differ between versions, and any other tool needs its own equivalent. + +A dataset of recorded multi-turn sessions, such as agent traces, must be +replayed by a tool that preserves each session's turn order and timing, since +they determine what the serving engine can reuse between requests. Sending the +dataset's requests independently does not benchmark that dataset. Such a +replay runs for a fixed duration rather than a number of requests. Use the same +replay options in every trial benchmark; the trial benchmarks may use a shorter +duration than the final benchmark, but all trials use the same one. A replay +may let a share of requests fail, up to the tool's failure threshold; keep the +tool's threshold and record the failed requests. Before any benchmark, ensure it uses a different seed than the previous benchmark. Otherwise the benchmark will depend on what has been cached by the @@ -373,7 +384,8 @@ works as expected: send real requests and check the responses, including reasoning output when the model supports it. These verification requests are never part of the measured metrics. -All verification and benchmark requests must succeed. +All verification and benchmark requests must succeed, except +for the failures a replay allows (see above). Record every benchmark using the following structure and field names — trial benchmarks in `trials//trial.json`, the final benchmark as @@ -400,10 +412,18 @@ trial benchmarks in `trials//trial.json`, the final benchmark as Set `workload.dataset` to `dataset` from `constraints.json`, and compute `workload.input_tokens` and `workload.output_tokens` as the measured mean input and output token counts of the benchmark, rounded to whole tokens. + +For a replay, set `duration_seconds` to the replay duration, not the tool's +total run time, which also includes the warmup and the wait for in-flight +requests to finish. Set `per_user_tok_per_s` to the tool's own mean per-user +output throughput: sessions idle between turns, so fewer than `concurrency` +requests are in flight on average. + Compute `output_tok_per_s` as `total_output_tokens / duration_seconds` and -`per_user_tok_per_s` as `output_tok_per_s / workload.concurrency`. These are -the numbers used to compare trials (see `## Performance`). +`per_user_tok_per_s` as `output_tok_per_s / workload.concurrency`, +except for a replay (see above). These are the numbers used to compare +trials (see `## Performance`). Set `tool` to the command name and subcommands without options or values, `tool_version` to the exact version, and `command` to the secret-free diff --git a/src/dstack/_internal/core/models/configurations.py b/src/dstack/_internal/core/models/configurations.py index 81d5d6226a..3ffc30b586 100644 --- a/src/dstack/_internal/core/models/configurations.py +++ b/src/dstack/_internal/core/models/configurations.py @@ -1814,7 +1814,8 @@ class PresetConfiguration( Field( description=( "The benchmark dataset used during preset creation: a benchmark tool's" - " dataset name (e.g. `sharegpt`, `spec_bench`) or a Hugging Face dataset ID." + " dataset name (e.g. `sharegpt`, `spec_bench`, or an AIPerf public dataset" + " such as `semianalysis_cc_traces_weka_062126_256k`) or a Hugging Face dataset ID." " Omit for synthetic prompts shaped by `input_tokens` and `output_tokens`" ) ), diff --git a/src/tests/_internal/cli/services/presets/test_prompt.py b/src/tests/_internal/cli/services/presets/test_prompt.py index a66092f02d..0327e0da4e 100644 --- a/src/tests/_internal/cli/services/presets/test_prompt.py +++ b/src/tests/_internal/cli/services/presets/test_prompt.py @@ -70,6 +70,8 @@ def test_renders_only_the_custom_dataset_branch(self): ) assert "`workload.dataset`" in text + assert "`aiperf profile`" in text + assert "For a replay, set `duration_seconds`" in text # The request shape is the random dataset's contract, not this one's. assert "shared_prefix_tokens" not in text @@ -80,6 +82,8 @@ def test_a_random_dataset_session_never_hears_of_datasets(self): assert "shared_prefix_tokens" in text assert "`dataset`" not in text + assert "replay" not in text + assert "All verification and benchmark requests must succeed." in text def test_fails_loudly_when_the_prompt_has_no_directives(self, tmp_path, monkeypatch): plain = tmp_path / "system_prompt.md" From 5b0454d0efed1d5b0adc96d22e17f295c5e26f0c Mon Sep 17 00:00:00 2001 From: Victor Skvortsov Date: Wed, 7 Oct 2026 17:13:10 +0500 Subject: [PATCH 2/2] Keep public docs unchanged for trace-replay datasets --- mkdocs/docs/concepts/presets.md | 13 ------------- src/dstack/_internal/core/models/configurations.py | 3 +-- 2 files changed, 1 insertion(+), 15 deletions(-) diff --git a/mkdocs/docs/concepts/presets.md b/mkdocs/docs/concepts/presets.md index 8a2b12e905..5a0ef995da 100644 --- a/mkdocs/docs/concepts/presets.md +++ b/mkdocs/docs/concepts/presets.md @@ -216,19 +216,6 @@ The requests every benchmark measures. The dataset provides the requests, so `input_tokens`, `output_tokens`, and `shared_prefix_tokens` can't be set with it, and the preset records the measured means. A gated dataset requires `HF_TOKEN` in `env`. - `dataset` can also name a dataset of recorded multi-turn sessions, such as one of the public agentic trace datasets of [SemiAnalysis's AIPerf fork](https://github.com/SemiAnalysisAI/aiperf). The agent then replays the sessions with their turn order and timing, so that prefix reuse matches real traffic: - - ```yaml - dataset: semianalysis_cc_traces_weka_062126_256k - concurrency: 16 - - prompt: | - Benchmark with AIPerf `--scenario inferencex-agentx-mvp`. Use a 600-second - replay for trials and a 3600-second replay for the final benchmark. - ``` - - A replay runs for a fixed duration rather than a number of requests, so creation takes much longer than with synthetic prompts. Use `prompt` to set the replay options, such as the duration. - ### Baseline By default, the first trial is a baseline: the agent serves the model the way the chosen serving framework recommends, without tuning it for performance. Later trials are optimization attempts. Set `baseline: false` to make every trial an optimization attempt. diff --git a/src/dstack/_internal/core/models/configurations.py b/src/dstack/_internal/core/models/configurations.py index 3ffc30b586..81d5d6226a 100644 --- a/src/dstack/_internal/core/models/configurations.py +++ b/src/dstack/_internal/core/models/configurations.py @@ -1814,8 +1814,7 @@ class PresetConfiguration( Field( description=( "The benchmark dataset used during preset creation: a benchmark tool's" - " dataset name (e.g. `sharegpt`, `spec_bench`, or an AIPerf public dataset" - " such as `semianalysis_cc_traces_weka_062126_256k`) or a Hugging Face dataset ID." + " dataset name (e.g. `sharegpt`, `spec_bench`) or a Hugging Face dataset ID." " Omit for synthetic prompts shaped by `input_tokens` and `output_tokens`" ) ),