Feat/qarnot cloud - #34
Open
florian-simvia wants to merge 42 commits into
Open
florian-simvia wants to merge 42 commits into
florian-simvia wants to merge 42 commits into
Conversation
Five verbs the core programs against, so a remote execution service can be plugged in without the runner learning its API. BackendState already speaks csauto's status vocabulary; translating a provider's own names is the backend's job.
What StubAdapter is to solvers: it satisfies the contract with local files and a scripted state sequence, so the remote lifecycle can be tested with no network, no account and no credit. It can also be told to fail a poll or a final download, which the sync pass has to survive.
A third branch beside the local process and the Slurm job. The case carries a backend name and a task id instead of a PID, and a failed submission leaves it FAILED rather than PENDING.
A backend case has neither a PID nor a scheduler job, so should_finalize was true on the first refresh and the STATUS_FAILED fallback marked it failed before the remote service had started it.
Polls each unfinished backend case once, appends the output deltas to csauto.stdout and csauto.stderr (which the Log Tail and Recent Errors already read), pulls the targeted snapshot, and marks the case DONE only once the results are actually on local disk. A failed poll never changes the status. Three phases like _count_running_cases: snapshot under a brief lock, network work with no lock held, write back under a new one.
Deliberately not inside refresh_status, which runs about twice a second because the dashboard polls every second and the server caches for half of one. A network round trip there would abuse the provider's API and make the dashboard lag. Two cadences: a light poll, and a less frequent file sync. The test drives TestClient as a context manager, because the shared helper does not and the FastAPI lifespan never starts there.
Submit, observe, finish on the fake backend, with no network. Documents the execution backend boundary and its two invariants.
The _backend throttle marker sits at the top of registry.json beside the cases, and refresh_status listed it as a phantom case in PREPARED. Registry iteration now goes through case_records.
A remote backend cannot call find_residuals_files on a directory that only exists on the compute node, so the adapter declares the patterns instead. The argv submitted to a backend now uses a relative case path for the same reason.
Bucket naming, the glob to regex translation Qarnot's snapshot filter needs, the per-case upload plan and its cost guard, and the shared-directory fingerprint. No SDK, no network, no account.
Profile, snapshot interval and per-case upload limit. A token in the config file is rejected outright: that file is shared, committed and archived, so the credential belongs in QARNOT_TOKEN.
Two input buckets per campaign: the shared mesh once, the case inputs per case. The snapshot whitelist comes from the adapter's observability globs, so only a few hundred kB come back while the run is in progress.
The state table lives in the backend, so runner.py never sees a Qarnot name, and an unknown state maps to PENDING rather than failing a live campaign. fetch_final pulls the results bucket unconditionally, because DONE means the results are on local disk.
…kend doctor --backend qarnot checks the SDK, the token, the image and whether the configured solver can build a remote command, each separately, and never prints the token itself.
The dashboard asks the server what is on offer rather than hardcoding a provider name, the same way it already does for panels and capabilities.
/api/run_case accepts a backend name, validated against the registered ones, and the status row carries back what the backend reports: progress, execution time and core count.
The execution list comes from /api/app_config, so the dialog never names a provider. Choosing one shows what is about to be submitted and billed before the button is pressed.
Execution time and core count come from the backend, exist for every solver and live in the registry, so they belong beside Duration rather than inside the solver-declared Timing panel.
Asserts the three claims the design rests on and cannot be checked against a fake: the inputs arrive in the working directory, a snapshot comes back mid run, and the full results land on local disk.
Qarnot refuses to abort a finished task with 'Invalid operation on non-running task'. Pressing Stop on a case that completed between the last refresh and the click returned HTTP 500. Found by the live round trip against a real account.
An exhausted bucket or storage quota surfaced as a QuotaExceeded in the middle of an upload, after a campaign had been chosen and launched. doctor now says what the account has room for, before anything is spent, and the live test stops leaking a bucket set per run.
Submitting uploads the case and talks to a remote API. Running that inside registry_transaction blocked every reader for its whole duration, so the dashboard froze on /api/status while the case sat in PENDING. The submit now runs unlocked and writes its outcome back under a short transaction.
REGISTRY_THREAD_LOCK is an RLock, so a nested call passed through it, but flock applies per descriptor and the nested call opened a second one, which waited on the first. The result was a silent, permanent freeze of the whole process. Launching on the Qarnot backend hit it before reaching the provider.
The rank and thread counts only reached the solver, inside DOCKER_CMD; the task itself carried no hardware constraint, so Qarnot allocated any available machine and MPI could oversubscribe on compute the user pays for. Submission now demands at least ranks x threads cores.
A backend says what a user may choose; the core renders the list and hands the chosen values back untouched, so a provider vocabulary stays in its module.
The runner and the route transport the map without reading it, and record it in the case history so an audit shows what a run asked for.
Declared as pure data so the module stays testable without the optional SDK, with drift guards comparing them to the SDK own classes when it is present.
Both setters raise once a task is launched, so they run before submit, and an unknown scheduling value fails the launch with a message naming the choices.
launch_options never raises: the dialog opens and a launch stays possible whatever the provider is doing, and a degraded result is not cached so a brief outage does not hide the node list for ten minutes.
A dedicated route, called when the dialog opens rather than on every page load, which degrades instead of failing so a launch is always possible.
The dialog renders whatever the backend declared, so it names no provider and no option key, and it opens even when the provider cannot be reached. Also writes the core-count documentation and changelog entry that a blocked command dropped when that fix was committed.
Qarnot validates constraints against a per-account catalogue and refuses the whole submission for one it does not recognise, so the core constraint csauto synthesises broke every launch with "Some constraints don t exist". It is now kept only when the catalogue lists it. The node type comes from that same catalogue, so it is valid by construction and always travels.
The case inputs were uploaded at the root of the remote working directory, so the shared directories landed inside the case instead of beside it, and code_saturne failed in preprocessing looking for its mesh in the study dir. The remote layout now mirrors the local one. The results a task returns are restricted to its own directory, so a download cannot reach RUNS/MESH, which is a symlink to the user own mesh.
A remote task has one writable directory, so a backend puts the campaign shared directories inside the case rather than beside it. code_saturne resolves a bare mesh name against the study directory, which does not exist there, and failed in preprocessing. Adapters gain prepare_remote_case, called before any backend launch; code_saturne adds a meshdir entry to setup.xml, verified locally to fix the remote layout and to change nothing locally. A task also returns only its results directory now, so the shared mesh it was handed is not downloaded back into the case on every run.
While a run is in progress the points come from the solver log, and each was numbered by counting printed convergence blocks. Convergence is printed at the listing frequency, so the Nth block is almost never iteration N, and a run at step 5000 showed a curve ending at 44. The step the solver announces is used instead, with the block count kept as a fallback.
Refresh reread local files, which for a case an execution backend owns only change when the background sync pass runs, so the panels looked frozen between two passes. A new throttled route pulls first; csauto status shares the same throttle, which moves next to the pass itself. The RESU size was cached against the results directory mtime, and some filesystems never bump a directory mtime when files change inside it, so the figure stayed frozen for the life of the process. The entry now expires.
The case selector was rendered inside the same condition as the plot, so selecting a case that had not run hid both and left no way to pick another. The controls now render whenever the campaign has cases. The Probes and Profiles tabs had a second cause: the loader overwrote the flag that enables a tab, which answers whether the campaign has any probe data, with one scoped to the current selection. That disabled the tab itself.
Neither earned its place: Qarnot only reports an execution time once a task has finished, so the column stayed empty for the whole run, and the core count showed zero for every local case. The figures are still recorded per case in the registry, where they cost nothing and remain available to an audit.
code_saturne removes its run_status.running marker when a calculation ends, but a copy pulled back by a backend snapshot survives, and read_progress kept preferring it, so a finished case reported the iteration the snapshot caught. Progress now comes from the log once a case is terminal. A second defect surfaced with it: extract_last_iteration did not recognise TIME STEP NUMBER, the marker the solver announces, and only worked when an incidental warning happened to mention a step.
A run killed by the runaway-computation check was reported DONE: code_saturne prints its closing banner before aborting and writes the real message to error rather than the log, so the log read as a success while the explicit run_status.failed marker beside it went unread. The marker now decides and the log is the fallback; the existing recency guard still keeps a marker from an earlier run out of a good one.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Run a campaign on the Qarnot cloud
Adds a second boundary to csauto, execution, alongside the solver boundary it already had, and implements it for Qarnot. A user can send cases to their own Qarnot account from the dashboard and watch them exactly as if they were running locally: residuals, probes, log tail and error panels all work live, with no cloud-specific code anywhere in the UI.
Verified end to end on a real Qarnot account: a code_saturne case submitted from the dashboard, observed while running, and its results downloaded.
What a user gets
Flex/OnDemand/Reserved) and Node type, read from the user's own account.csauto doctor --backend qarnot, reporting each prerequisite separately: SDK, token, image, whether the solver can build a remote command, and what the account still has room for.The Qarnot token is read from
QARNOT_TOKENand rejected if found incsauto.toml: that file is shared, committed and archived. No route, payload or error message ever carries it.Architecture
ExecutionBackend(csauto/backends/base.py) is the new contract: five verbs (submit,poll,sync,fetch_final,cancel), pluslaunch_options()for what a backend lets a user choose.BackendState.statusalready speaks csauto's vocabulary, so translating a provider's own state names is the backend's job, never the runner's.Everything Qarnot-specific lives in
csauto/backends/qarnot.py; the parts that are decisions rather than API calls live inqarnot_support.pyand are unit-tested with no SDK installed. The optional[qarnot]extra is imported inside the backend module only, exactly as FastAPI is today.Two invariants the implementation depends on:
DONEmeans the results are on local disk. Compare reads files, Clean removes directories,resu_size_mbmeasures a size. A case is markedDONEonly afterfetch_finalsucceeded.The boundary holds in both directions:
grep -rni "qarnot|OnDemand|scheduling"returns nothing underfrontend/src, andrunner.pyand the routes carry the launch options without interpreting them. The frontend does not even know the option keys — it renders what/api/launch_optionsreturns.localand Slurm are deliberately not ported to the contract. They predate it and work; porting them is a later, mechanical refactor.Adapter additions
Two declarations, both consumed by generic code:
observability_globs— which files a backend should pull back while a run is in progress. Needed becausefind_residuals_filesand friends inspect a local directory and cannot describe a task running on someone else's machine.prepare_remote_case(case_dir)— a remote task has one writable directory and no parent, so a backend puts the campaign's shared directories inside the case. A solver that expects them elsewhere is adjusted by its own adapter. code_saturne adds a<meshdir>entry tosetup.xml; the entry is harmless locally, where the study directory stays a fallback.Defects found and fixed along the way
Most were found by running against a real account, and none would have shown up against a fake.
registry_lockre-entrancyREGISTRY_THREAD_LOCKis anRLock, so a nested call passed through it, butflockapplies per descriptor and the nested call opened a second one, which waited on the first. A silent, permanent freeze of the whole process. Launching on Qarnot hit it before a single request reached the provider.registry_transactionblocked every reader for its duration, so the dashboard froze on/api/status._backendthrottle marker sat beside the cases andrefresh_statuslisted it as a phantom case.LAST ITERon a finished caserun_status.runningmarker pulled back by a snapshot survives the end of a run, and progress kept being read from it. Also,extract_last_iterationdid not recogniseTIME STEP NUMBER, the marker the solver actually announces.error, so the log heuristic read a runaway computation as a success. The explicitrun_status.failedmarker now decides.RESUsize frozenTesting
ruff check,ruff format --checkandprettierclean.FakeBackend(backend = "fake") covers the whole remote lifecycle with no network and no account, the way thestubsolver does without a solver.tests/integration/test_qarnot_live.pyis the one test that costs money: skipped unlessCSAUTO_QARNOT_LIVE=1and aQARNOT_TOKENare set. It asserts the three claims the design rests on and cannot be checked against a fake: the case inputs arrive in the task's working directory, a partial result comes back while the task is still running, and the full results land on local disk. It has been run against a real account and passes.importorskipguard compares the two whenever the SDK is installed, so a rename upstream is caught.qarnotSDK uninstalled;doctor --backend qarnotreports the missing extra instead of raising.Not in this PR
RESU/makes that tree a read-only resource directory, and the solver can no longer create its new run directory. Making it work means placing the checkpoint outsideRESUand pointing the solver at it without breaking local runs. Restart still works locally, unchanged.doctorsays so instead of submitting something that will fail.instancecount > 1, and deleting a campaign's buckets when it is finished.Known residue
A case relaunched locally keeps its
backendfield in the registry, so it carries both apidand a backend name. Harmless today but contradictory, and worth cleaning up before local and Slurm move behind the contract.