Skip to content

Feat/qarnot cloud - #34

Open
florian-simvia wants to merge 42 commits into
mainfrom
feat/qarnot-cloud
Open

florian-simvia wants to merge 42 commits into
mainfrom
feat/qarnot-cloud

Conversation

@florian-simvia

Copy link
Copy Markdown
Collaborator

Run a campaign on the Qarnot cloud

Adds a second boundary to csauto, execution, alongside the solver boundary it already had, and implements it for Qarnot. A user can send cases to their own Qarnot account from the dashboard and watch them exactly as if they were running locally: residuals, probes, log tail and error panels all work live, with no cloud-specific code anywhere in the UI.

Verified end to end on a real Qarnot account: a code_saturne case submitted from the dashboard, observed while running, and its results downloaded.

What a user gets

  • A Run on selector in the Run dialog, defaulting to This machine. The choice is per launch, not per campaign, so one case can be verified locally while the other ninety-nine go to the cloud.
  • Per-launch Priority (Flex / OnDemand / Reserved) and Node type, read from the user's own account.
  • A confirmation restating how many cases are about to be submitted and that they are billed to the user's own account.
  • Live monitoring during a cloud run. Only the files the solver adapter declares come back while the run is in progress, a few hundred kB, so the dashboard is fed without downloading a multi-gigabyte results directory.
  • csauto doctor --backend qarnot, reporting each prerequisite separately: SDK, token, image, whether the solver can build a remote command, and what the account still has room for.

The Qarnot token is read from QARNOT_TOKEN and rejected if found in csauto.toml: that file is shared, committed and archived. No route, payload or error message ever carries it.

Architecture

ExecutionBackend (csauto/backends/base.py) is the new contract: five verbs (submit, poll, sync, fetch_final, cancel), plus launch_options() for what a backend lets a user choose. BackendState.status already speaks csauto's vocabulary, so translating a provider's own state names is the backend's job, never the runner's.

Everything Qarnot-specific lives in csauto/backends/qarnot.py; the parts that are decisions rather than API calls live in qarnot_support.py and are unit-tested with no SDK installed. The optional [qarnot] extra is imported inside the backend module only, exactly as FastAPI is today.

Two invariants the implementation depends on:

  • DONE means the results are on local disk. Compare reads files, Clean removes directories, resu_size_mb measures a size. A case is marked DONE only after fetch_final succeeded.
  • A failed poll never changes a status. A flaky network must not mark a hundred cases as failed; consecutive failures are counted instead.

The boundary holds in both directions: grep -rni "qarnot|OnDemand|scheduling" returns nothing under frontend/src, and runner.py and the routes carry the launch options without interpreting them. The frontend does not even know the option keys — it renders what /api/launch_options returns.

local and Slurm are deliberately not ported to the contract. They predate it and work; porting them is a later, mechanical refactor.

Adapter additions

Two declarations, both consumed by generic code:

  • observability_globs — which files a backend should pull back while a run is in progress. Needed because find_residuals_files and friends inspect a local directory and cannot describe a task running on someone else's machine.
  • prepare_remote_case(case_dir) — a remote task has one writable directory and no parent, so a backend puts the campaign's shared directories inside the case. A solver that expects them elsewhere is adjusted by its own adapter. code_saturne adds a <meshdir> entry to setup.xml; the entry is harmless locally, where the study directory stays a fallback.

Defects found and fixed along the way

Most were found by running against a real account, and none would have shown up against a fake.

Fix What it was
registry_lock re-entrancy REGISTRY_THREAD_LOCK is an RLock, so a nested call passed through it, but flock applies per descriptor and the nested call opened a second one, which waited on the first. A silent, permanent freeze of the whole process. Launching on Qarnot hit it before a single request reached the provider.
Registry lock across a submit Submitting uploads a case and talks to a remote API; running that inside registry_transaction blocked every reader for its duration, so the dashboard froze on /api/status.
Hardware constraints Qarnot validates constraints against a per-account catalogue and refuses a whole submission for one it does not recognise. A synthesised core constraint broke every launch.
Remote case layout The shared directories landed inside the case rather than beside it, and code_saturne failed in preprocessing looking for its mesh in the study directory.
Reserved registry key The _backend throttle marker sat beside the cases and refresh_status listed it as a phantom case.
Live residual abscissa Points were numbered by counting printed convergence blocks. Convergence is printed at the listing frequency, so a run at step 5000 showed a curve ending at 44. Affects local runs too.
LAST ITER on a finished case A run_status.running marker pulled back by a snapshot survives the end of a run, and progress kept being read from it. Also, extract_last_iteration did not recognise TIME STEP NUMBER, the marker the solver actually announces.
Outcome of a diverged run code_saturne prints its closing banner before aborting and writes the real message to error, so the log heuristic read a runaway computation as a success. The explicit run_status.failed marker now decides.
Stopping a finished task Qarnot refuses to abort a finished task, so pressing Stop on a case that completed between two refreshes returned HTTP 500.
RESU size frozen The cached figure was keyed on the results directory's mtime, and some filesystems never bump a directory's mtime when files change inside it, so the cache never invalidated.
Plot panels unusable Selecting a case with no results hid the case selector along with the plot, leaving no way to pick another. In Probes and Profiles the loader also overwrote the flag that enables a tab, disabling the tab itself.

Testing

  • 551 passed, 4 skipped. ruff check, ruff format --check and prettier clean.
  • FakeBackend (backend = "fake") covers the whole remote lifecycle with no network and no account, the way the stub solver does without a solver.
  • tests/integration/test_qarnot_live.py is the one test that costs money: skipped unless CSAUTO_QARNOT_LIVE=1 and a QARNOT_TOKEN are set. It asserts the three claims the design rests on and cannot be checked against a fake: the case inputs arrive in the task's working directory, a partial result comes back while the task is still running, and the full results land on local disk. It has been run against a real account and passes.
  • Where a local object stands in for an SDK class, an importorskip guard compares the two whenever the SDK is installed, so a rename upstream is caught.
  • csauto still imports and runs with the qarnot SDK uninstalled; doctor --backend qarnot reports the missing extra instead of raising.

Not in this PR

  • Restart on an execution backend. It was implemented and then reverted: attaching the previous run's checkpoint under RESU/ makes that tree a read-only resource directory, and the solver can no longer create its new run directory. Making it work means placing the checkpoint outside RESU and pointing the solver at it without breaking local runs. Restart still works locally, unchanged.
  • The Stop control on a cloud case, which very likely meets the same read-only constraint: the control file has to land in a directory the solver creates at runtime.
  • code_aster on Qarnot. Its adapter composes its whole launch line locally and produces no argv a remote container can run; doctor says so instead of submitting something that will fail.
  • Porting local and Slurm to the contract, memory and CPU-model launch options, per-campaign defaults, instancecount > 1, and deleting a campaign's buckets when it is finished.

Known residue

A case relaunched locally keeps its backend field in the registry, so it carries both a pid and a backend name. Harmless today but contradictory, and worth cleaning up before local and Slurm move behind the contract.

Five verbs the core programs against, so a remote execution service can be
plugged in without the runner learning its API. BackendState already speaks
csauto's status vocabulary; translating a provider's own names is the backend's
job.
What StubAdapter is to solvers: it satisfies the contract with local files and a
scripted state sequence, so the remote lifecycle can be tested with no network,
no account and no credit. It can also be told to fail a poll or a final
download, which the sync pass has to survive.
A third branch beside the local process and the Slurm job. The case carries a
backend name and a task id instead of a PID, and a failed submission leaves it
FAILED rather than PENDING.
A backend case has neither a PID nor a scheduler job, so should_finalize was
true on the first refresh and the STATUS_FAILED fallback marked it failed before
the remote service had started it.
Polls each unfinished backend case once, appends the output deltas to
csauto.stdout and csauto.stderr (which the Log Tail and Recent Errors already
read), pulls the targeted snapshot, and marks the case DONE only once the
results are actually on local disk. A failed poll never changes the status.

Three phases like _count_running_cases: snapshot under a brief lock, network
work with no lock held, write back under a new one.
Deliberately not inside refresh_status, which runs about twice a second because
the dashboard polls every second and the server caches for half of one. A
network round trip there would abuse the provider's API and make the dashboard
lag. Two cadences: a light poll, and a less frequent file sync.

The test drives TestClient as a context manager, because the shared helper does
not and the FastAPI lifespan never starts there.
Submit, observe, finish on the fake backend, with no network. Documents the
execution backend boundary and its two invariants.
The _backend throttle marker sits at the top of registry.json beside the
cases, and refresh_status listed it as a phantom case in PREPARED. Registry
iteration now goes through case_records.
A remote backend cannot call find_residuals_files on a directory that only
exists on the compute node, so the adapter declares the patterns instead. The
argv submitted to a backend now uses a relative case path for the same reason.
Bucket naming, the glob to regex translation Qarnot's snapshot filter needs,
the per-case upload plan and its cost guard, and the shared-directory
fingerprint. No SDK, no network, no account.
Profile, snapshot interval and per-case upload limit. A token in the config
file is rejected outright: that file is shared, committed and archived, so the
credential belongs in QARNOT_TOKEN.
Two input buckets per campaign: the shared mesh once, the case inputs per
case. The snapshot whitelist comes from the adapter's observability globs, so
only a few hundred kB come back while the run is in progress.
The state table lives in the backend, so runner.py never sees a Qarnot name,
and an unknown state maps to PENDING rather than failing a live campaign.
fetch_final pulls the results bucket unconditionally, because DONE means the
results are on local disk.
…kend

doctor --backend qarnot checks the SDK, the token, the image and whether the
configured solver can build a remote command, each separately, and never
prints the token itself.
The dashboard asks the server what is on offer rather than hardcoding a
provider name, the same way it already does for panels and capabilities.
/api/run_case accepts a backend name, validated against the registered ones,
and the status row carries back what the backend reports: progress, execution
time and core count.
The execution list comes from /api/app_config, so the dialog never names a
provider. Choosing one shows what is about to be submitted and billed before
the button is pressed.
Execution time and core count come from the backend, exist for every solver
and live in the registry, so they belong beside Duration rather than inside
the solver-declared Timing panel.
Asserts the three claims the design rests on and cannot be checked against a
fake: the inputs arrive in the working directory, a snapshot comes back mid
run, and the full results land on local disk.
Qarnot refuses to abort a finished task with 'Invalid operation on non-running
task'. Pressing Stop on a case that completed between the last refresh and the
click returned HTTP 500. Found by the live round trip against a real account.
An exhausted bucket or storage quota surfaced as a QuotaExceeded in the middle
of an upload, after a campaign had been chosen and launched. doctor now says
what the account has room for, before anything is spent, and the live test
stops leaking a bucket set per run.
Submitting uploads the case and talks to a remote API. Running that inside
registry_transaction blocked every reader for its whole duration, so the
dashboard froze on /api/status while the case sat in PENDING. The submit now
runs unlocked and writes its outcome back under a short transaction.
REGISTRY_THREAD_LOCK is an RLock, so a nested call passed through it, but
flock applies per descriptor and the nested call opened a second one, which
waited on the first. The result was a silent, permanent freeze of the whole
process. Launching on the Qarnot backend hit it before reaching the provider.
The rank and thread counts only reached the solver, inside DOCKER_CMD; the
task itself carried no hardware constraint, so Qarnot allocated any available
machine and MPI could oversubscribe on compute the user pays for. Submission
now demands at least ranks x threads cores.
A backend says what a user may choose; the core renders the list and hands the
chosen values back untouched, so a provider vocabulary stays in its module.
The runner and the route transport the map without reading it, and record it
in the case history so an audit shows what a run asked for.
Declared as pure data so the module stays testable without the optional SDK,
with drift guards comparing them to the SDK own classes when it is present.
Both setters raise once a task is launched, so they run before submit, and an
unknown scheduling value fails the launch with a message naming the choices.
launch_options never raises: the dialog opens and a launch stays possible
whatever the provider is doing, and a degraded result is not cached so a brief
outage does not hide the node list for ten minutes.
A dedicated route, called when the dialog opens rather than on every page
load, which degrades instead of failing so a launch is always possible.
The dialog renders whatever the backend declared, so it names no provider and
no option key, and it opens even when the provider cannot be reached.

Also writes the core-count documentation and changelog entry that a blocked
command dropped when that fix was committed.
Qarnot validates constraints against a per-account catalogue and refuses the
whole submission for one it does not recognise, so the core constraint csauto
synthesises broke every launch with "Some constraints don t exist". It is now
kept only when the catalogue lists it. The node type comes from that same
catalogue, so it is valid by construction and always travels.
The case inputs were uploaded at the root of the remote working directory, so
the shared directories landed inside the case instead of beside it, and
code_saturne failed in preprocessing looking for its mesh in the study dir.
The remote layout now mirrors the local one. The results a task returns are
restricted to its own directory, so a download cannot reach RUNS/MESH, which
is a symlink to the user own mesh.
A remote task has one writable directory, so a backend puts the campaign
shared directories inside the case rather than beside it. code_saturne
resolves a bare mesh name against the study directory, which does not exist
there, and failed in preprocessing. Adapters gain prepare_remote_case, called
before any backend launch; code_saturne adds a meshdir entry to setup.xml,
verified locally to fix the remote layout and to change nothing locally.

A task also returns only its results directory now, so the shared mesh it was
handed is not downloaded back into the case on every run.
While a run is in progress the points come from the solver log, and each was
numbered by counting printed convergence blocks. Convergence is printed at the
listing frequency, so the Nth block is almost never iteration N, and a run at
step 5000 showed a curve ending at 44. The step the solver announces is used
instead, with the block count kept as a fallback.
Refresh reread local files, which for a case an execution backend owns only
change when the background sync pass runs, so the panels looked frozen between
two passes. A new throttled route pulls first; csauto status shares the same
throttle, which moves next to the pass itself.

The RESU size was cached against the results directory mtime, and some
filesystems never bump a directory mtime when files change inside it, so the
figure stayed frozen for the life of the process. The entry now expires.
The case selector was rendered inside the same condition as the plot, so
selecting a case that had not run hid both and left no way to pick another.
The controls now render whenever the campaign has cases.

The Probes and Profiles tabs had a second cause: the loader overwrote the flag
that enables a tab, which answers whether the campaign has any probe data,
with one scoped to the current selection. That disabled the tab itself.
Neither earned its place: Qarnot only reports an execution time once a task
has finished, so the column stayed empty for the whole run, and the core count
showed zero for every local case. The figures are still recorded per case in
the registry, where they cost nothing and remain available to an audit.
code_saturne removes its run_status.running marker when a calculation ends,
but a copy pulled back by a backend snapshot survives, and read_progress kept
preferring it, so a finished case reported the iteration the snapshot caught.
Progress now comes from the log once a case is terminal.

A second defect surfaced with it: extract_last_iteration did not recognise
TIME STEP NUMBER, the marker the solver announces, and only worked when an
incidental warning happened to mention a step.
A run killed by the runaway-computation check was reported DONE: code_saturne
prints its closing banner before aborting and writes the real message to
error rather than the log, so the log read as a success while the explicit
run_status.failed marker beside it went unread. The marker now decides and the
log is the fallback; the existing recency guard still keeps a marker from an
earlier run out of a good one.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant