diff --git a/src/dstack/_internal/cli/services/presets/resources/system_prompt.md b/src/dstack/_internal/cli/services/presets/resources/system_prompt.md index 691df34b33..d367b8f799 100644 --- a/src/dstack/_internal/cli/services/presets/resources/system_prompt.md +++ b/src/dstack/_internal/cli/services/presets/resources/system_prompt.md @@ -342,10 +342,21 @@ example, the dataset options are: | --- | --- | | `vllm bench serve` | `--dataset-name ` when `dataset` is the tool's own dataset name, or `--dataset-name hf --dataset-path ` when it is a Hugging Face dataset ID | | `sglang.benchmark.serving` | `--dataset-name ` when `dataset` is the tool's own dataset name; the tool has no Hugging Face dataset option | +| `aiperf profile` | `--public-dataset ` when `dataset` is one of AIPerf's public datasets; for agentic trace datasets such as `semianalysis_cc_traces_weka_062126_256k`, also `--scenario inferencex-agentx-mvp` (SemiAnalysis's AIPerf fork), which replays multi-turn sessions with their timing | The table is an example and not a full command: the remaining options still come from `concurrency`, option names and defaults differ between versions, and any other tool needs its own equivalent. + +A dataset of recorded multi-turn sessions, such as agent traces, must be +replayed by a tool that preserves each session's turn order and timing, since +they determine what the serving engine can reuse between requests. Sending the +dataset's requests independently does not benchmark that dataset. Such a +replay runs for a fixed duration rather than a number of requests. Use the same +replay options in every trial benchmark; the trial benchmarks may use a shorter +duration than the final benchmark, but all trials use the same one. A replay +may let a share of requests fail, up to the tool's failure threshold; keep the +tool's threshold and record the failed requests. Before any benchmark, ensure it uses a different seed than the previous benchmark. Otherwise the benchmark will depend on what has been cached by the @@ -373,7 +384,8 @@ works as expected: send real requests and check the responses, including reasoning output when the model supports it. These verification requests are never part of the measured metrics. -All verification and benchmark requests must succeed. +All verification and benchmark requests must succeed, except +for the failures a replay allows (see above). Record every benchmark using the following structure and field names — trial benchmarks in `trials//trial.json`, the final benchmark as @@ -400,10 +412,18 @@ trial benchmarks in `trials//trial.json`, the final benchmark as Set `workload.dataset` to `dataset` from `constraints.json`, and compute `workload.input_tokens` and `workload.output_tokens` as the measured mean input and output token counts of the benchmark, rounded to whole tokens. + +For a replay, set `duration_seconds` to the replay duration, not the tool's +total run time, which also includes the warmup and the wait for in-flight +requests to finish. Set `per_user_tok_per_s` to the tool's own mean per-user +output throughput: sessions idle between turns, so fewer than `concurrency` +requests are in flight on average. + Compute `output_tok_per_s` as `total_output_tokens / duration_seconds` and -`per_user_tok_per_s` as `output_tok_per_s / workload.concurrency`. These are -the numbers used to compare trials (see `## Performance`). +`per_user_tok_per_s` as `output_tok_per_s / workload.concurrency`, +except for a replay (see above). These are the numbers used to compare +trials (see `## Performance`). Set `tool` to the command name and subcommands without options or values, `tool_version` to the exact version, and `command` to the secret-free diff --git a/src/tests/_internal/cli/services/presets/test_prompt.py b/src/tests/_internal/cli/services/presets/test_prompt.py index a66092f02d..0327e0da4e 100644 --- a/src/tests/_internal/cli/services/presets/test_prompt.py +++ b/src/tests/_internal/cli/services/presets/test_prompt.py @@ -70,6 +70,8 @@ def test_renders_only_the_custom_dataset_branch(self): ) assert "`workload.dataset`" in text + assert "`aiperf profile`" in text + assert "For a replay, set `duration_seconds`" in text # The request shape is the random dataset's contract, not this one's. assert "shared_prefix_tokens" not in text @@ -80,6 +82,8 @@ def test_a_random_dataset_session_never_hears_of_datasets(self): assert "shared_prefix_tokens" in text assert "`dataset`" not in text + assert "replay" not in text + assert "All verification and benchmark requests must succeed." in text def test_fails_loudly_when_the_prompt_has_no_directives(self, tmp_path, monkeypatch): plain = tmp_path / "system_prompt.md"