This project is a solution to the DataHow Machine Learning Engineer coding challenge. The goal is twofold:
- Part 1 — Develop a predictive model for the final product titer (mg/L) of a simulated upstream bioprocess for monoclonal antibody (mAb) production.
- Part 2 — Implement an inference microservice exposing a REST API that serves predictions from the trained model.
The dataset is a panel of bioprocess experiments. Each experiment runs a bioreactor under a specific recipe and records time-series measurements of controlled inputs and biological state variables. The prediction target is a single scalar per experiment: the final antibody concentration at harvest.
Each row in the CSV files corresponds to one time point of one experiment. Variables are grouped by their role in the process:
| Variable | Description |
|---|---|
Z:FeedStart |
Day on which glucose and glutamine feeding begins |
Z:FeedEnd |
Day on which feeding stops (typically the harvest day) |
Z:FeedRateGlc |
Glucose feed rate (g/L/day) |
Z:FeedRateGln |
Glutamine feed rate (g/L/day) |
Z:phStart |
Initial pH setpoint at experiment start |
Z:phEnd |
pH setpoint after the pH shift (production phase) |
Z:phShift |
Day on which the pH setpoint changes |
Z:tempStart |
Initial temperature setpoint (°C) during growth phase |
Z:tempEnd |
Temperature setpoint (°C) after the temperature shift |
Z:tempShift |
Day on which the temperature setpoint changes |
Z:Stir |
Stirring rate (rpm) |
Z:DO |
Dissolved oxygen setpoint (%) |
Z:ExpDuration |
Planned experiment duration in days (harvest day) |
Z variables represent the experimental design — the recipe set before the run
starts. They are recorded only at t=0; all subsequent rows carry NaN by design.
| Variable | Description |
|---|---|
W:temp |
Measured temperature (°C) at each time point |
W:pH |
Measured pH at each time point |
W:FeedGlc |
Glucose feed applied (g/L) — zero before Z:FeedStart |
W:FeedGln |
Glutamine feed applied (g/L) — zero before Z:FeedStart |
W variables are the actuator setpoints applied during the run. They follow the recipe but may deviate slightly due to process noise and control dynamics.
| Variable | Description |
|---|---|
X:VCD |
Viable Cell Density (10⁶ cells/mL) — primary growth metric |
X:Glc |
Glucose concentration (g/L) |
X:Gln |
Glutamine concentration (g/L) |
X:Amm |
Ammonia concentration (mM) — metabolic by-product |
X:Lac |
Lactate concentration (g/L) — metabolic by-product |
X:Lysed |
Lysed (dead) cell density (10⁶ cells/mL) |
X variables are the biological response to the W inputs and Z recipe. They are measured at each sampling time point throughout the experiment.
| Variable | Description |
|---|---|
Y:Titer |
Final mAb product concentration (mg/L) at harvest |
flowchart LR
Z["**Z** — Recipe parameters<br/>static, recorded at t=0<br/>13 variables"] -->|"sets process conditions"| W["**W** — Controlled inputs<br/>time-varying setpoints<br/>4 variables"]
Z -->|"fixes experiment duration"| Y["**Y** — Titer<br/>final harvest outcome<br/>mg/L"]
W -->|"drives biological response"| X["**X** — Biological state<br/>measured time series<br/>6 variables"]
X -->|"cumulative cell dynamics"| Y
.
├── data/
│ ├── datahow_interview_train_data.csv
│ ├── datahow_interview_train_targets.csv
│ ├── datahow_interview_test_data.csv
│ └── datahow_interview_test_targets-TEMPLATE.csv
├── src/
│ ├── data_loading.py # I/O, column constants, structural transforms
│ ├── eda.py # EDA plots and statistics
│ ├── features.py # feature matrix construction
│ ├── model.py # TiterPredictor class
│ ├── evaluation.py # cross-validation and plots
│ ├── train.py # training orchestrator
│ ├── predict.py # test-set inference and evaluation
│ └── server/ # Part 2: inference microservice
│ ├── app.py # FastAPI app with lifespan and endpoints
│ ├── schemas.py # Pydantic v2 request/response DTOs
│ └── inference.py # payload → raw_df → features → prediction
├── output/
│ ├── eda/ # EDA figures (python src/eda.py)
│ ├── test_predictions.csv # test-set predictions (python src/predict.py)
│ └── models/
│ ├── ridge/ # artefacts for the Ridge model
│ │ ├── model.joblib
│ │ ├── actual_vs_predicted.png
│ │ ├── residuals.png
│ │ ├── log_residuals.png
│ │ └── coefficients.png
│ ├── elasticnet/ # artefacts for the ElasticNet model
│ │ ├── model.joblib
│ │ ├── actual_vs_predicted.png
│ │ ├── residuals.png
│ │ ├── log_residuals.png
│ │ └── coefficients.png
│ ├── rf/ # artefacts for the Random Forest model
│ │ ├── model.joblib
│ │ ├── actual_vs_predicted.png
│ │ ├── residuals.png
│ │ ├── log_residuals.png
│ │ └── feature_importance.png
│ └── ridge_adapted/ # artefacts for the Ridge adapted model
│ ├── model.joblib
│ ├── actual_vs_predicted.png
│ ├── residuals.png
│ ├── log_residuals.png
│ └── coefficients.png
├── tests/
│ ├── conftest.py # shared fixtures (synthetic predictor, TestClient)
│ └── test_server.py # integration tests for all endpoints
├── examples/
│ ├── predict_request.json # sample POST /predict payload
│ └── predict_response.json # expected response for the sample payload
├── config.yaml # pipeline configuration (paths, active model)
├── requirements.txt # ML pipeline dependencies
├── requirements-server.txt # inference server dependencies
├── Dockerfile # container image for the inference server
├── pyproject.toml # Black + Ruff + pytest configuration
└── README.md
pip install -r requirements.txtEDA
python src/eda.pySaves figures to output/eda/.
Training
python src/train.pyTrains Ridge, ElasticNet, Random Forest, and the adapted Ridge variant, runs
5-fold stratified CV for all models, prints a comparison table, and saves
artefacts to output/models/{ridge,elasticnet,rf,ridge_adapted}/.
Test-set inference
train.py automatically sets model_path in config.yaml to the model with
the lowest CV RMSE. To use a specific model instead, edit config.yaml manually:
# config.yaml — set model_path to the desired model:
# model_path: output/models/ridge/model.joblib
# model_path: output/models/elasticnet/model.joblib
# model_path: output/models/rf/model.joblib
# model_path: output/models/ridge_adapted/model.joblib # distribution-shift variantUpdate targets_path to point to the test targets CSV, then run:
python src/predict.pySaves output/test_predictions.csv with columns Y:Titer_actual and
Y:Titer_predicted. Evaluation metrics (RMSE, NRMSE, MAE, MAPE, R²) are printed
to stdout and an actual-vs-predicted plot is saved to
output/models/{model_type}/actual_vs_predicted_test.png.
Z columns are not missing data. Z variables encode the experimental recipe —
set once before the run starts and recorded only at t=0. All subsequent rows
carry NaN by design. extract_z_features filters to t=0 to extract the 13
scalar values cleanly.
Everything is aggregated to experiment level. The target is one scalar per
experiment, but the input is a time series of varying length. Z variables are
already scalars; W and X variables are aggregated using three statistics per
variable: the last observed value (_last), the temporal mean (_mean), and the
area under the curve (_auc, trapezoid rule). All three are computed in the EDA
to enable direct comparison.
_last value features are excluded. Training experiments run for 7, 8, 9,
10, or 14 days; all test experiments are exactly 14 days. A last-value feature
captures the measurement at the final time point, which occurs at different
biological stages
across experiments of different lengths — making the values incomparable between
train and test. Mean and AUC aggregate over the full trajectory and are robust to
this duration shift.
IVC is included as an engineered feature. IVC (Integral of Viable Cell Density) is the area under the X:VCD curve over time, computed with the trapezoid rule (unit: 10⁶ cells·day/mL). In mAb bioprocessing, antibody titer is classically proportional to the cumulative viable cell exposure: more cell-hours in the bioreactor means more antibody produced. IVC is a domain-standard feature that compresses the entire VCD time series into one interpretable scalar.
log(Y:Titer) is used for modelling. Raw titer has a right-skewed distribution (skewness ≈ +1.88). The log-transformed target is near-symmetric (skewness ≈ −0.055), which stabilises variance and benefits linear models. The transform is applied only at model training and evaluation — it is not stored in the data tables.
Each insight is traceable to a specific output of python src/eda.py.
| Insight | Console section | Plot |
|---|---|---|
| log(Y:Titer) is near-Gaussian (skewness −0.055 vs +1.88 raw) | Target Distribution (Y:Titer) |
eda_target_distribution.png |
| All 20 test experiments run exactly 14 days; train spans 7–14 | Experiment Duration Distribution |
eda_duration_distribution.png |
| Z:ExpDuration matches the actual time-series length for all 120 experiments (normalization by planned duration is exact) | Planned vs Actual Experiment Duration |
— |
| Z:ExpDuration is the strongest Z predictor (r ≈ +0.62) | Z Variables — Pearson Correlation with Y:Titer |
eda_z_pearson.png |
| X:Lysed dominates aggregated feature correlations (mean r ≈ +0.69, AUC r ≈ +0.65) | Aggregated W/X Features — Pearson Correlation with Y:Titer |
eda_aggregated_correlations.png |
_last features are excluded due to duration shift — they capture measurements at different biological stages across experiments of different lengths |
Experiment Duration Distribution |
— |
| IVC correlates strongly with titer | IVC vs Y:Titer |
eda_ivc_vs_titer.png |
| Q4 experiments separate from Q1 early in the run on X:VCD and X:Lysed | — | eda_x_timeseries.png |
The full feature matrix is built from three sources in src/features.py:
- Z features (13) — scalars extracted from
t=0rows, one per experiment, representing the experimental recipe. - W/X aggregations (19) — mean and AUC per variable (4 W + 6 X = 10 variables × 2,
minus
X:VCD_auc)._lastis excluded due to the train/test duration shift (see 1.1).X:VCD_aucis excluded because it is mathematically identical to IVC (both compute ∫ X:VCD dt via the trapezoid rule); keeping both would introduce a redundant feature with opposite-sign coefficients — a collinearity artefact. IVC is retained under its domain name. - IVC (1) — ∫ X:VCD dt per experiment (trapezoid rule).
AUC features are normalised by Z:ExpDuration (divided by the planned duration
in days). Raw AUC grows proportionally with experiment length regardless of biology:
a 14-day experiment accumulates roughly twice the AUC of a 7-day one even if the
underlying dynamics are identical. This duration confounding is critical here because
all test experiments run for exactly 14 days while the training set spans 7–14 days
(durations: 7, 8, 9, 10, and 14 days; 90% shorter than 14 days). Dividing by
duration converts AUC into a time-averaged
rate, comparable across all experiment lengths.
Total: 33 features per experiment.
Model: regularised regression in log space. TiterPredictor wraps a
StandardScaler + regressor pipeline. Training is performed on log(Y:Titer)
(motivated by the near-Gaussian log-target from the EDA); predictions are
back-transformed with exp() before scoring.
Why Ridge over ordinary least squares (OLS). With 100 experiments and 33 features, plain OLS tends
to overfit and produce unstable coefficients. L2 regularisation (Ridge) reduces
this by adding a penalty that shrinks coefficients toward zero, trading a small
amount of bias for lower variance. The strength of the penalty is controlled by
alpha: a larger alpha shrinks coefficients more aggressively (higher bias,
lower variance); a smaller alpha approaches OLS (lower bias, higher variance).
RidgeCV selects the optimal alpha from the explicit grid
[0.01, 0.1, 1.0, 3.0, 5.0, 10.0, 30.0, 100.0, 1000.0] via 5-fold
cross-validation. Selected alpha=3.0.
Why ElasticNet as a second model. ElasticNet combines L1 (sparsity) and L2
(stability) regularisation. The L1 term drives irrelevant feature coefficients
exactly to zero, performing automatic feature selection. ElasticNetCV selects
both alpha and l1_ratio via 5-fold CV over the extended grids
[0.0001, 0.001, 0.005, 0.01, 0.05, 0.1, 0.5, 1.0, 5.0, 10.0] and
[0.1, 0.5, 0.7, 0.9, 0.95, 1.0]. Selected alpha=0.005, l1_ratio=1.0
(pure LASSO) — zeros out 15 of 33 features and retains 18 informative predictors.
Why Random Forest as a third model. Random Forest is the natural
non-parametric comparison point. It makes no linearity assumptions, captures
feature interactions automatically, and provides Mean Decrease in Impurity (MDI) feature importances
directly from the fitted trees. RandomForestRegressor(n_estimators=500, random_state=42) is wrapped in a GridSearchCV using the shared _INNER_CV
splitter for consistency with Ridge and ElasticNet. The grid searches
max_depth ∈ [None, 5, 10, 20], min_samples_leaf ∈ [1, 2, 5, 10], and
max_features ∈ ["sqrt", 0.3, 0.5, 0.7, 1.0] — the hyperparameters most
sensitive to overfitting on small datasets. With 100 training samples and 33
features, ensemble methods are typically disadvantaged relative to regularised
linear models; the RF serves to confirm whether non-linear interactions provide
additional signal beyond what linear models can capture. For this reason, RF is
not used as a base for the adapted variant. Best hyperparameters selected by
inner CV: {'max_depth': 10, 'max_features': 0.5, 'min_samples_leaf': 1}.
Adapted model (distribution shift mitigation). All test experiments run
exactly 14 days, while only ~10% of the training set does. This creates a
distribution shift with two consequences: Z:ExpDuration is constant on the
test set (zero discriminative power there) and the model has seen few 14-day
experiments across the full titer range. The adapted variant addresses this with
two changes applied on top of the best linear base model (Ridge):
Z:ExpDurationis removed from the input features. It is still used internally inbuild_feature_matrixto normalise AUC and IVC features, but not passed to the regressor.- Inverse-frequency instance weighting: each experiment is weighted by
1 / freq(duration), normalised so the mean weight equals 1. This automatically upweights under-represented duration groups (14-day experiments) with no manually tuned multiplier.
The outer CV of the adapted model uses the same titer-stratified folds for a fair comparison, but its CV RMSE is not directly comparable to the base models: the validation folds contain mixed durations while the adapted model is optimised for 14-day experiments. A lower CV RMSE for a base model does not imply better performance on the all-14-day test set.
Evaluation: nested cross-validation (random_state=42, fully reproducible).
Training and hyperparameter selection are kept strictly separate through two
independent CV loops.
The outer loop (StratifiedKFold(5), stratified by titer quartile) splits the
100 training experiments into 5 folds of 80/20. Each held-out fold of 20
experiments is never seen during training or hyperparameter selection — it is used
only to score the model. Averaging the 5 held-out scores gives an unbiased estimate
of generalisation performance. Stratification by titer quartile ensures each fold
contains a representative mix of low- and high-titer experiments, reducing
fold-to-fold variance (plain KFold produced R² std ±0.299 due to titer clustering;
stratification reduces this to ±0.068).
The inner loop operates on each fold's 80-experiment training set. RidgeCV
tries every alpha in the explicit grid; ElasticNetCV searches over all
(alpha, l1_ratio) combinations; RF uses GridSearchCV over depth, leaf size and
feature fraction. All inner CV uses the shared _INNER_CV splitter
(KFold(5, random_state=42)), so hyperparameter selection never touches the outer
held-out set. The fold-level models are discarded after scoring.
After CV, a fresh TiterPredictor is trained on all 100 experiments for each model
type. The inner CV selects the best hyperparameters on the full training set, and
this final model is serialised to disk. The CV score estimates generalisation
performance; the final model uses all available data for deployment. Metrics are computed
in the original mg/L scale after back-transforming predictions:
- RMSE / MAE (Root Mean Square Error / Mean Absolute Error) — absolute error in mg/L, directly interpretable as typical prediction error.
- NRMSE — RMSE normalised by the full target range (283–4822 mg/L), expressed as a percentage. Provides scale-free context: an NRMSE of 7 % means the typical error is 7 % of the observable target range.
- MAPE — mean absolute percentage error. Complements RMSE by expressing error relative to the actual titer of each experiment, useful when errors on low-titer runs are as important as errors on high-titer runs.
Diagnostic plots are saved in-sample per model to output/models/{model_type}/.
Linear models (ridge, elasticnet, ridge_adapted) produce
coefficients.png; the RF model produces feature_importance.png (MDI).
| Plot | Description |
|---|---|
actual_vs_predicted.png |
Scatter of predicted vs actual titer (mg/L) with perfect-prediction diagonal |
residuals.png |
Residuals (ŷ − y) vs actual titer in mg/L — checks for systematic bias |
log_residuals.png |
Residuals in log space (log ŷ − log y) vs log y — checks bias in the modelling scale |
coefficients.png |
Standardised coefficients (log mg/L per std dev), sorted by absolute value (linear models) |
feature_importance.png |
MDI feature importances sorted by value (RF only) |
| Metric | Ridge | ElasticNet | RF | Ridge_Adapted† |
|---|---|---|---|---|
| RMSE | 276.0 ± 95.3 mg/L | 298.0 ± 117.8 mg/L | 400.4 ± 105.3 mg/L | 308.5 ± 122.7 mg/L |
| NRMSE | 6.1 ± 2.1 % | 6.6 ± 2.6 % | 8.8 ± 2.3 % | 6.8 ± 2.7 % |
| MAE | 188.5 ± 33.5 mg/L | 197.0 ± 49.0 mg/L | 242.5 ± 44.1 mg/L | 211.5 ± 41.5 mg/L |
| MAPE | 14.9 ± 2.1 % | 14.8 ± 2.4 % | 17.9 ± 3.4 % | 16.7 ± 1.2 % |
| R² | 0.854 ± 0.068 | 0.829 ± 0.082 | 0.709 ± 0.086 | 0.822 ± 0.094 |
The RF result confirms that no non-linear interactions above what regularised linear models capture are detectable with 100 training samples. With more data, RF or gradient boosting would be the natural next step.
† CV RMSE for the adapted model is not directly comparable — see Design Decisions above.
Ridge is selected as the final model (lowest mean CV RMSE). Selected alpha=3.0.
ElasticNet selected alpha=0.005, l1_ratio=1.0 (pure LASSO — 15 of 33
features zeroed out, 18 retained).
Coefficient interpretation. Model coefficients are in log(mg/L) per
standard deviation of each feature: a coefficient c means that a one-std-dev
increase in the feature multiplies the predicted titer by exp(c). Ridge retains
all 33 features with non-zero but shrunken coefficients; ElasticNet (pure LASSO,
l1_ratio=1.0) explicitly zeroes 15 features and selects 18. The top-5 by
magnitude for each model:
| Feature | Ridge coef | ElasticNet coef | Direction |
|---|---|---|---|
X:Lac_mean |
+0.216 | +0.406 | positive — strongest signal in both |
X:Lac_auc |
+0.183 | ≈ 0 (zeroed) | collinear with X:Lac_mean; L1 drops it |
Z:ExpDuration |
+0.169 | +0.222 | positive — longer runs produce more |
X:Gln_mean (glutamine) |
−0.135 | −0.235 | negative — consistent across both models |
W:FeedGln_auc |
+0.101 | +0.136 | positive — consistent across both models |
X:Glc_mean (glucose) |
— | −0.201 | negative — not in Ridge top 5; ElasticNet explicitly selects it |
The directionality is fully consistent across both models. X:Lac_mean is the
dominant predictor in both: higher mean lactate reflects high metabolic activity
and antibody productivity. X:Lac_auc is dropped by ElasticNet (L1 selects the
more informative mean statistic when the two are collinear). X:Gln_mean and
X:Glc_mean are negative — high residual substrate concentrations suggest
suboptimal consumption and lower productivity. W:FeedGln_auc is positive in
both, reflecting that sustained glutamine feeding extends productive culture.
A FastAPI microservice that serves the trained model via two endpoints defined
in inference_server_spec.yml.
Install server dependencies (requires a trained model from Part 1):
pip install -r requirements-server.txtrequirements-server.txt is the unified runtime requirements file: it contains
the ML pipeline packages needed at inference time (pandas, numpy, scipy,
scikit-learn, joblib, pyyaml) together with the server framework
dependencies (fastapi, uvicorn, pydantic). requirements.txt retains the
full development dependencies (formatters, linters, visualisation libraries) and
is used by the devcontainer postCreateCommand.
Two options are available to start the server (reads model_path from
config.yaml). Option A is recommended for development; Option B runs the
exact production image.
Option A — local development
PYTHONPATH=src uvicorn server.app:app --host 0.0.0.0 --port 8080 --reloadOption B — Docker
docker build -t datahow-server .
docker run -p 8080:8080 datahow-serverBoth options expose the same two endpoints on port 8080:
GET /health— returns{"status": "ok"}POST /predict— accepts a JSON payload withtimestampsandvalues, returns predicted titer plus model metadata (model_type,model_version)
Tests (no running server required):
python -m pytest tests/ -vGET /health
Liveness check. Returns {"status": "ok"} with HTTP 200. No authentication or
model state is required.
POST /predict
Accepts a JSON payload describing one bioprocess experiment and returns the predicted final titer.
A full example request is provided in examples/predict_request.json
(14-day experiment from the OpenAPI spec, 15 time points spanning days 0–14). Test the running server with:
curl -X POST http://localhost:8080/predict \
-H "Content-Type: application/json" \
-d @examples/predict_request.jsonExpected response (examples/predict_response.json):
{
"titer_predicted_mg_L": 1993.8552,
"model_type": "ridge",
"model_version": "465a92ceee45"
}timestamps: strictly increasing array of time points in days (non-empty, lengthn). Non-consecutive spacing is supported — AUC features use the trapezoid rule over the actual time values.values: variable arrays keyed by prefix convention:Z:— scalar recipe parameters, single-element arraysW:andX:— time-series arrays of lengthn
All 23 variables used during training are required (Z_COLS, W_COLS,
X_COLS from data_loading.py). The implementation satisfies all spec-mandated
validation rules and adds further constraints to prevent silent feature
corruption. Invalid payloads are rejected with HTTP 422 (FastAPI/Pydantic
standard for schema and content validation failures; the OpenAPI spec names 400,
but 422 is semantically correct for well-formed requests that fail content
validation).
Rules required by the spec:
- all required Z, W, and X variables must be present
- Z variables must be single-element arrays
- W and X arrays must have length equal to
len(timestamps)
Additional constraints enforced by the implementation:
- unknown top-level fields are rejected (
extra="forbid") timestampsmust be non-empty and strictly increasing (no duplicates, no out-of-order values)Z:ExpDurationmust be greater than 0Z:ExpDurationmust equalmax(timestamps)— AUC features are normalised byZ:ExpDurationinbuild_feature_matrix; a mismatch would corrupt features silently at inference time
titer_predicted_mg_L: predicted final product titer in mg/Lmodel_type: model variant used ("ridge","ridge_adapted","elasticnet","rf")model_version: SHA-256 fingerprint (12 hex chars) of themodel.joblibfile loaded at startup — ties each prediction to the exact serialised artefact
POST /predict
↓
PredictRequest (Pydantic v2 validation)
↓
inference.predict_from_payload()
├── _payload_to_raw_df() — reshape JSON → raw DataFrame (one row per time point)
├── build_feature_matrix() — aggregate to experiment-level feature vector (33 features)
└── TiterPredictor.predict() — StandardScaler → Ridge → exp() → mg/L
↓
PredictResponse
The payload is converted to the same raw DataFrame format used during training,
so the entire feature engineering pipeline (build_feature_matrix) runs
unchanged. The exclude_duration flag is read from the serialised model, so
the correct feature set is used automatically regardless of which model variant
is configured in config.yaml.
The predictor is loaded once at startup via FastAPI's lifespan context manager
and injected into each request via a Depends dependency, enabling clean
override in tests without touching the lifespan.
Deserialization (JSON → Python) and serialization (Python → JSON) are handled transparently by Pydantic v2 and FastAPI:
- Incoming request: FastAPI passes the JSON body to Pydantic, which coerces
JSON numbers to
float, arrays tolist[float], and objects todict[str, list[float]]. Unknown top-level fields are rejected with HTTP 422 (extra="forbid"). The parsed model is immutable (frozen=True) — no code path can mutate it after construction. - Outgoing response: FastAPI calls
PredictResponse.model_dump()and serialises the result to JSON.PredictResponseis immutable (frozen=True). Amath.isfiniteguard inpredict_from_payloadrejects non-finite prediction values before the response is constructed, so they are caught by the existingpost_predictexception handler and returned as a clean HTTP 500.
29 integration tests cover both endpoints, all validation paths, and internal
error handling. The test suite does not require a trained model file: a
synthetic TiterPredictor is fitted in conftest.py and injected via
app.dependency_overrides. The synthetic training set consists of 20 copies of
the spec example experiment's feature vector with small Gaussian noise (σ=0.01)
added to each of the 33 features, paired with random uniform targets in
[500, 2000] — enough to produce a well-defined Ridge fit with the correct
feature schema, without requiring the real dataset.
The tests cover only Part 2 (inference server). Unit tests for the Part 1
pipeline (features.py, model.py) were not implemented within the challenge
time frame and are listed under Possible Extensions.
tests/test_server.py::TestHealth::test_returns_200
tests/test_server.py::TestHealth::test_body_contains_ok_status
tests/test_server.py::TestPredict::test_returns_200
tests/test_server.py::TestPredict::test_response_contains_titer_key
tests/test_server.py::TestPredict::test_prediction_is_positive_float
tests/test_server.py::TestPredict::test_prediction_is_deterministic
tests/test_server.py::TestPredict::test_response_contains_model_type
tests/test_server.py::TestPredict::test_model_type_is_non_empty_string
tests/test_server.py::TestPredict::test_model_type_matches_fixture
tests/test_server.py::TestPredict::test_response_contains_model_version
tests/test_server.py::TestPredict::test_model_version_is_non_empty_string
tests/test_server.py::TestPredict::test_model_version_matches_fixture
tests/test_server.py::TestPredictValidation::test_missing_z_variable_returns_422
tests/test_server.py::TestPredictValidation::test_missing_x_variable_returns_422
tests/test_server.py::TestPredictValidation::test_missing_w_variable_returns_422
tests/test_server.py::TestPredictValidation::test_empty_timestamps_returns_422
tests/test_server.py::TestPredictValidation::test_z_variable_multi_element_array_returns_422
tests/test_server.py::TestPredictValidation::test_x_variable_length_mismatch_returns_422
tests/test_server.py::TestPredictValidation::test_missing_timestamps_field_returns_422
tests/test_server.py::TestPredictValidation::test_missing_values_field_returns_422
tests/test_server.py::TestPredictValidation::test_empty_body_returns_422
tests/test_server.py::TestPredictValidation::test_non_increasing_timestamps_returns_422
tests/test_server.py::TestPredictValidation::test_duplicate_timestamps_returns_422
tests/test_server.py::TestPredictValidation::test_zero_exp_duration_returns_422
tests/test_server.py::TestPredictValidation::test_negative_exp_duration_returns_422
tests/test_server.py::TestPredictValidation::test_exp_duration_mismatch_returns_422
tests/test_server.py::TestPredictValidation::test_extra_top_level_field_returns_422
tests/test_server.py::TestPredictErrorPaths::test_prediction_failure_returns_500
tests/test_server.py::TestPredictErrorPaths::test_non_finite_prediction_returns_500
Unit tests for the Part 1 pipeline. The current test suite covers the
inference server only. The highest-value additions would be: (1) unit tests for
build_feature_matrix — verify feature count (33), column names, and AUC
normalisation by Z:ExpDuration; (2) unit tests for TiterPredictor — verify
the save/load roundtrip and that predict returns positive values after
back-transforming from log space.
Finer hyperparameter tuning. The current grids use explicit value arrays
chosen to cover the plausible range without excessive computation. A natural
next step is automated grid densification: RidgeCV accepts arbitrarily many
alphas at no extra cost (SVD — Singular Value Decomposition — based), so np.logspace(-3, 4, 100) would find
near-continuous optima. For ElasticNet and RF, RandomizedSearchCV over
continuous distributions (e.g. loguniform for alpha, randint for tree
depth) would explore a much larger space within a fixed compute budget.
MLP on engineered features. The simplest non-linear extension is a shallow
MLP (Multi-Layer Perceptron) trained directly on the 33 features. ElasticNet's feature selection (18
non-zero out of 33) could serve as a preprocessing step, reducing dimensionality
before the MLP to limit overfitting. With 100 training samples overfitting risk is
high; sklearn's MLPRegressor offers only L2 regularisation, while a PyTorch
implementation would enable dropout and early stopping, making this approach more
viable even on small datasets.
1D CNN (Convolutional Neural Network) on raw time series. A 1D convolutional network could process the W and X time series directly, learning temporal features without manual aggregation. It would require a hybrid architecture (CNN branch for W/X sequences + dense branch for Z scalars), padding or masking to handle variable-length inputs (7–14 days), and PyTorch as an additional dependency. With 100 samples the risk of overfitting is substantial; this extension becomes practical with a larger dataset of bioprocess experiments.
I acted as architect and domain decision-maker throughout. All structural choices — what to build, in what order, and what to include or exclude — originated from me. Claude Code's role was to implement those decisions, surface technical constraints (API changes, type errors, style violations), and propose alternatives when asked — never to set direction autonomously.
Phase 1 — Problem understanding
"Analyze the challenge PDF and the dataset. Explain the problem, the variable groups, and what biological meaning each prefix (Z/W/X/Y) carries."
I built a shared understanding of the domain before committing to any design choice. This grounded all downstream decisions in the actual data structure and biological context. The Variable Interaction Overview Mermaid diagram was proposed and produced in this phase.
Phase 2 — EDA pipeline
"Start with an EDA. I want: target distribution, Pearson correlation of each variable with titer, and time-series profiles of W and X grouped by titer quartile."
I defined the scope of the analysis: correlation plots, quartile-stratified time-series profiles, histogram and Q-Q plots for the target.
Phase 3 — Aggregation strategy
"Aggregate everything to experiment level. For Z variables use only the t=0 rows — those are the recipe parameters. For W and X, reduce the time series to scalars using mean and AUC. Skip the last-value — it is not comparable across experiments of different durations."
I identified the train/test duration-shift problem and decided the aggregation scheme before implementation started.
Phase 4 — Domain knowledge and feature engineering
"What domain-specific features or tricks are standard for this type of bioprocess data? What does the literature typically use to predict antibody titer from cell culture time series?"
This led to the introduction of IVC (Integral of Viable Cell Density) — the trapezoid integral of X:VCD over time — as a domain-standard engineered feature that captures cumulative cell exposure and is classically proportional to final antibody titer.
Phase 5 — Model selection and evaluation framework
"Add ElasticNet as a second model type. Design a nested CV framework: outer StratifiedKFold(5) on titer quartile for generalisation estimation, inner 5-fold CV for hyperparameter selection (RidgeCV, ElasticNetCV). After CV, train a separate final model on all 100 experiments. Print a side-by-side comparison and automatically update config.yaml to point to the model with the lower RMSE."
Deciding which models to compare first, then designing the evaluation framework around them, keeps the CV design tightly scoped. The nested structure separates generalisation estimation (outer) from hyperparameter selection (inner): inner CV operates only on each fold's training set and never touches the held-out data. The saved model is trained from scratch on all 100 experiments.
Phase 6 — Test-set inference and configuration
"Write a predict.py script for test-set inference. Read all paths from a config.yaml. Save predictions with Y:Titer_actual and Y:Titer_predicted columns side by side. Print evaluation metrics to stdout only — not in the CSV."
This produced predict.py + config.yaml as the inference entry point. Updating
targets_path in config.yaml to the received test targets file and rerunning is
the only step needed to evaluate against real test targets.
Phase 7 — Distribution shift and adapted model
"The test set is entirely made of 14-day experiments while only ~10% of the training set does. Implement an adapted model variant: exclude Z:ExpDuration from the input features and add inverse-frequency instance weighting to upweight 14-day experiments."
The EDA confirmed that Z:ExpDuration is the strongest Z predictor in the training
set (r ≈ +0.62), but it is constant on the test set — providing no discriminative
signal there. The adapted variant addresses this with two changes on top of the best
linear base model: (1) Z:ExpDuration is removed from the input features (still
used internally for AUC normalisation); (2) inverse-frequency instance weighting
upweights under-represented duration groups automatically. The exclude_duration
flag is serialised inside the model so that inference code reconstructs the correct
feature matrix without additional configuration.
Phase 8 — Random Forest and hyperparameter grid refinement
"Add Random Forest as a third base model using GridSearchCV with the shared inner CV splitter. RF should not be used as a base for the adapted model. Add MDI feature importance plots. Also extend the hyperparameter grids for Ridge and ElasticNet."
Refining the grids changed the winner: Ridge improved from 303.3 to 276.0 mg/L
by selecting alpha=3.0 — a value not reachable with the original 6-point grid.
ElasticNet shifted to a smaller alpha (0.005) with pure LASSO (l1_ratio=1.0) —
more aggressive feature selection but higher CV RMSE than Ridge. RF underperforms
on this 100-sample dataset (RMSE 400.4 mg/L) even after tuning max_depth,
confirming that linear models are more appropriate at this data scale.
Phase 9 — Response schema design
"Design the POST /predict response schema. Keep the flat columnar request format that mirrors the CSV Z:/W:/X: prefix convention. Enrich the response with model_type and model_version: model_type derived from predictor.model_type and the exclude_duration flag; model_version as the first 12 hex chars of the SHA-256 of model.joblib."
model_type emits "ridge_adapted" when exclude_duration=True.
model_version is the first 12 hex chars of the SHA-256 hash of the serialised
model.joblib file, tying each prediction to the exact trained artefact.
Phase 10 — Example payload files
"Add an example input payload and an example response as local files in an
examples/directory."
examples/predict_request.json uses the spec example experiment (15 timestamps,
all 23 variables from inference_server_spec.yml).
examples/predict_response.json records the real Ridge output
(1993.8552 mg/L, version 465a92ceee45) so that the expected response is
traceable to the exact trained artefact.
Phase 11 — Additional input validations
"Add additional input validations to
PredictRequest: enforce strictly increasing timestamps,Z:ExpDuration > 0, andZ:ExpDuration == max(timestamps)to prevent silent feature corruption at inference time."
The decision was to enforce Z:ExpDuration == max(timestamps) at the API
boundary rather than changing the normalisation denominator in
build_feature_matrix (which would require retraining). Three validators were
added to PredictRequest in src/server/schemas.py: (1) timestamps strictly
increasing, checked in the timestamps_non_empty field validator; (2)
Z:ExpDuration > 0; (3) Z:ExpDuration == max(timestamps) within 1e-6 —
both in validate_values. Silently wrong predictions are replaced by explicit
422 responses. This phase produced the first 26 integration tests.
Phase 12 — DTO (Data Transfer Object) serialisation hardening
"Harden the Pydantic DTOs: add
extra='forbid'andfrozen=TruetoPredictRequest,frozen=TruetoPredictResponse, and amath.isfiniteguard inpredict_from_payloadto catch non-finite prediction values before the response is constructed."
model_config = ConfigDict(extra="forbid", frozen=True) added to
PredictRequest: unknown top-level fields are rejected with HTTP 422 and the
parsed model is immutable after construction. model_config = ConfigDict(frozen=True)
added to PredictResponse. A math.isfinite guard in predict_from_payload
catches non-finite prediction values before the response is constructed and
returns a clean HTTP 500 via the existing exception handler. One test added
for extra="forbid" brings the total to 27.
Phase 13 — Dockerization improvements
"Containerise the inference server:
python:3.13-slimbase image, a unifiedrequirements-server.txtthat includes both ML runtime and FastAPI/uvicorn/pydantic dependencies, and aHEALTHCHECKusingurllib.request."
requirements-server.txt was rewritten as the unified runtime requirements file,
merging ML packages with the server framework and removing dev-only dependencies
from the container. devcontainer.json postCreateCommand updated to install
both requirements.txt and requirements-server.txt explicitly.