Skip to content

Add presets user interview - #4337

Open
r4victor wants to merge 1 commit into
masterfrom
pr_preset_interview
Open

r4victor wants to merge 1 commit into
masterfrom
pr_preset_interview

Conversation

@r4victor

@r4victor r4victor commented Oct 6, 2026

Copy link
Copy Markdown
Collaborator

This PR extends dstack-presets skill with a user interview process that allows collecting requirements from the user in a structured way. It allows writing presets configurations using agents effectively and predictably.

Here's the skill in action:

❯ I want to optimize Qwen-3.8-Flash-Next inference on 4xH200

You asked to optimize Qwen/Qwen3.8-Flash-Next inference on 4xH200. In dstack that means creating a preset, and the presets skill on this branch says to interview you first. Before I write preset.dstack.yml, I need a few answers from you.

What I've confirmed so far

- Model: Qwen/Qwen3.8-Flash-Next is an MoE with 180B parameters, 512 experts and 10 active per token. Its max context is 262144. It isn't gated, so no HF_TOKEN is needed. There is an official -FP8 variant.
- Memory: it fits on 4xH200 (564 GB total).
  - BF16 weights take about 360 GB; FP8 takes about 180 GB.
  - KV cache is small: only 12 of the 48 layers use full attention, with 2 KV heads, so it's about 24 KB per token. A full 256K-token request needs about 6.4 GB.
- Hardware: no fleet is limited to H200. cloud-fleet has no resource limits. It does return 4xH200 offers (runpod, jarvislabs, vastai, from $18.36/h), but the agent could also pick other GPUs from it.
- Earlier presets: none exist for this model. The qwen3-27b-agentic* presets are for Qwen3.8-27B, so I won't pass them as previous.

Questions (my proposed value is in brackets)

1. Model variant: is FP8 acceptable for your quality bar? If yes, I'll use base, which lets the agent choose FP8 or BF16. If quality must not change, I'll pin the exact weights with repo. [base]
2. Fleet: the preset should run only on the hardware it's optimized for. Should I create an h200-fleet with gpu: H200:4, on-demand, nodes: 0..1? Tell me if you want to limit it to certain backends or set a max price. [any backend, no price cap]
3. Workload: what does a typical request look like? Pick a profile below and adjust it, or point me to a dataset:
   - Chat: 2000 in / 500 out, 1000 shared prefix
   - RAG: 8000 in / 500 out, 1000 shared prefix
   - Coding agent: 32000 in / 1000 out, 28000 shared prefix
   - Summarization: 8000 in / 1000 out, no shared prefix
   - Reasoning: 1000 in / 8000 out, no shared prefix
   - Your 27B presets used 128K in / 512 out with a 99% shared prefix. Should this one match them?
4. Concurrency: how many requests run at the same time per replica? Your earlier presets used 4 and 8. If you only know requests per second or user counts, give me those and I'll work it out.
5. TTFT limit: what is the maximum p50 time to first token, in ms? I'll suggest a value once I know the workload.
6. Context length: what is the longest request the endpoint must accept? [262144, the model maximum, which is what your earlier presets used]
7. Trials: how many benchmarked trials should the agent run? [4]
8. Optional: do you want a specific agent model or effort level? Should the baseline trial be skipped? Is there anything the agent must try or must not change, such as a framework (vLLM or SGLang) or source patching?

Once you answer, I'll show you the full YAML. It will list what presets can't enforce, such as p99 latency, cost, or required serving features like tool calling. I'll run dstack apply only after you accept it.

needs input: answers to questions 1-7 above (FP8 OK?, create h200-fleet?, workload shape, concurrency, max TTFT, context length, trials)

✻ Baked for 1m 27s · done 3:30 PM · 1 shell still running

❯ 1. fp8 is fine
  2. yes, create h200-fleet in jarvislabs
  3. Coding agent: 32000 in / 1000 out, 28000 shared prefix
  4. concurrency 4-8
  5. suggest ttft
  6. 262144
  7. 4

You said FP8 is fine, to create h200-fleet on jarvislabs, and to use the coding-agent workload (32000 in / 1000 out, 28000 shared prefix), concurrency 4-8, a context length of 262144, and 4 trials. You asked me to suggest the TTFT limit. I'll create the fleet now, check the preset field names in the code, then write the YAML for you to confirm.


- No HF token is needed (the model isn't gated).
- The agent uses your default claude CLI settings.
- A baseline trial runs first.
- No earlier presets are passed as previous; this is the first preset for this model.

What the preset can't enforce

- The agent only maximizes output tokens per second at concurrency 8.
- The only hard limits are p50 TTFT and context length. Nothing checks:
  - p90/p99 TTFT, time per output token, or speed per user;
  - performance at concurrency 4;
  - cost;
  - output quality after FP8 quantization;
  - serving features a coding agent needs, such as tool calling, tool-call parsing, or reasoning output.
- If you need tool calling or a specific framework (vLLM or SGLang), I can add it to prompt, but the agent treats that as guidance, not a check.

needs input: Do you accept this configuration (concurrency 8, max_ttft 1000 ms), or do you want changes, such as two presets for concurrency 4 and 8, or a tool-calling note in prompt?

✻ Churned for 53s · done 3:34 PM · 1 shell still running

❯ Confirm

Preset creation is running as qwen38-flash-next-coding-h200 (ID cea1a59b). It's on trial 1 of 4, and the only hardware it has is the 4xH200 jarvislabs offer at $18.36/h.

What the agent has decided so far:
- It picked Qwen/Qwen3.8-Flash-Next-FP8, the official block-quantized checkpoint (about 173 GiB). That leaves about 95 GB per GPU at TP4 for KV cache and linear-attention state.
- It ruled out BF16 because the weights take twice the memory, which slows decoding. It ruled out NVFP4 because H200 has no FP4 tensor cores.
- It asks for a 500 GB disk, since the default 100 GB can't hold the checkpoint.

I'll keep watching and report when it finishes. You can also follow it with dstack preset logs -f cea1a59b, or stop it with dstack preset stop cea1a59b.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant