Skip to content

Repository files navigation

DataHow ML Engineer Coding Challenge

Challenge Overview

This project is a solution to the DataHow Machine Learning Engineer coding challenge. The goal is twofold:

  1. Part 1 — Develop a predictive model for the final product titer (mg/L) of a simulated upstream bioprocess for monoclonal antibody (mAb) production.
  2. Part 2 — Implement an inference microservice exposing a REST API that serves predictions from the trained model.

The dataset is a panel of bioprocess experiments. Each experiment runs a bioreactor under a specific recipe and records time-series measurements of controlled inputs and biological state variables. The prediction target is a single scalar per experiment: the final antibody concentration at harvest.


Dataset and Variables

Each row in the CSV files corresponds to one time point of one experiment. Variables are grouped by their role in the process:

Z — Static recipe parameters (recorded once at t=0)

Variable Description
Z:FeedStart Day on which glucose and glutamine feeding begins
Z:FeedEnd Day on which feeding stops (typically the harvest day)
Z:FeedRateGlc Glucose feed rate (g/L/day)
Z:FeedRateGln Glutamine feed rate (g/L/day)
Z:phStart Initial pH setpoint at experiment start
Z:phEnd pH setpoint after the pH shift (production phase)
Z:phShift Day on which the pH setpoint changes
Z:tempStart Initial temperature setpoint (°C) during growth phase
Z:tempEnd Temperature setpoint (°C) after the temperature shift
Z:tempShift Day on which the temperature setpoint changes
Z:Stir Stirring rate (rpm)
Z:DO Dissolved oxygen setpoint (%)
Z:ExpDuration Planned experiment duration in days (harvest day)

Z variables represent the experimental design — the recipe set before the run starts. They are recorded only at t=0; all subsequent rows carry NaN by design.

W — Controlled time-varying inputs

Variable Description
W:temp Measured temperature (°C) at each time point
W:pH Measured pH at each time point
W:FeedGlc Glucose feed applied (g/L) — zero before Z:FeedStart
W:FeedGln Glutamine feed applied (g/L) — zero before Z:FeedStart

W variables are the actuator setpoints applied during the run. They follow the recipe but may deviate slightly due to process noise and control dynamics.

X — Measured biological state variables

Variable Description
X:VCD Viable Cell Density (10⁶ cells/mL) — primary growth metric
X:Glc Glucose concentration (g/L)
X:Gln Glutamine concentration (g/L)
X:Amm Ammonia concentration (mM) — metabolic by-product
X:Lac Lactate concentration (g/L) — metabolic by-product
X:Lysed Lysed (dead) cell density (10⁶ cells/mL)

X variables are the biological response to the W inputs and Z recipe. They are measured at each sampling time point throughout the experiment.

Y — Target

Variable Description
Y:Titer Final mAb product concentration (mg/L) at harvest

Variable Interaction Overview

flowchart LR
    Z["**Z** — Recipe parameters<br/>static, recorded at t=0<br/>13 variables"] -->|"sets process conditions"| W["**W** — Controlled inputs<br/>time-varying setpoints<br/>4 variables"]
    Z -->|"fixes experiment duration"| Y["**Y** — Titer<br/>final harvest outcome<br/>mg/L"]
    W -->|"drives biological response"| X["**X** — Biological state<br/>measured time series<br/>6 variables"]
    X -->|"cumulative cell dynamics"| Y
Loading

Code Structure

.
├── data/
│   ├── datahow_interview_train_data.csv
│   ├── datahow_interview_train_targets.csv
│   ├── datahow_interview_test_data.csv
│   └── datahow_interview_test_targets-TEMPLATE.csv
├── src/
│   ├── data_loading.py     # I/O, column constants, structural transforms
│   ├── eda.py              # EDA plots and statistics
│   ├── features.py         # feature matrix construction
│   ├── model.py            # TiterPredictor class
│   ├── evaluation.py       # cross-validation and plots
│   ├── train.py            # training orchestrator
│   ├── predict.py          # test-set inference and evaluation
│   └── server/             # Part 2: inference microservice
│       ├── app.py          # FastAPI app with lifespan and endpoints
│       ├── schemas.py      # Pydantic v2 request/response DTOs
│       └── inference.py    # payload → raw_df → features → prediction
├── output/
│   ├── eda/                # EDA figures (python src/eda.py)
│   ├── test_predictions.csv  # test-set predictions (python src/predict.py)
│   └── models/
│       ├── ridge/              # artefacts for the Ridge model
│       │   ├── model.joblib
│       │   ├── actual_vs_predicted.png
│       │   ├── residuals.png
│       │   ├── log_residuals.png
│       │   └── coefficients.png
│       ├── elasticnet/         # artefacts for the ElasticNet model
│       │   ├── model.joblib
│       │   ├── actual_vs_predicted.png
│       │   ├── residuals.png
│       │   ├── log_residuals.png
│       │   └── coefficients.png
│       ├── rf/                 # artefacts for the Random Forest model
│       │   ├── model.joblib
│       │   ├── actual_vs_predicted.png
│       │   ├── residuals.png
│       │   ├── log_residuals.png
│       │   └── feature_importance.png
│       └── ridge_adapted/      # artefacts for the Ridge adapted model
│           ├── model.joblib
│           ├── actual_vs_predicted.png
│           ├── residuals.png
│           ├── log_residuals.png
│           └── coefficients.png
├── tests/
│   ├── conftest.py         # shared fixtures (synthetic predictor, TestClient)
│   └── test_server.py      # integration tests for all endpoints
├── examples/
│   ├── predict_request.json  # sample POST /predict payload
│   └── predict_response.json # expected response for the sample payload
├── config.yaml             # pipeline configuration (paths, active model)
├── requirements.txt        # ML pipeline dependencies
├── requirements-server.txt # inference server dependencies
├── Dockerfile              # container image for the inference server
├── pyproject.toml          # Black + Ruff + pytest configuration
└── README.md

Part 1: Predictive Model

Setup

pip install -r requirements.txt

Running

EDA

python src/eda.py

Saves figures to output/eda/.

Training

python src/train.py

Trains Ridge, ElasticNet, Random Forest, and the adapted Ridge variant, runs 5-fold stratified CV for all models, prints a comparison table, and saves artefacts to output/models/{ridge,elasticnet,rf,ridge_adapted}/.

Test-set inference

train.py automatically sets model_path in config.yaml to the model with the lowest CV RMSE. To use a specific model instead, edit config.yaml manually:

# config.yaml — set model_path to the desired model:
# model_path: output/models/ridge/model.joblib
# model_path: output/models/elasticnet/model.joblib
# model_path: output/models/rf/model.joblib
# model_path: output/models/ridge_adapted/model.joblib  # distribution-shift variant

Update targets_path to point to the test targets CSV, then run:

python src/predict.py

Saves output/test_predictions.csv with columns Y:Titer_actual and Y:Titer_predicted. Evaluation metrics (RMSE, NRMSE, MAE, MAPE, R²) are printed to stdout and an actual-vs-predicted plot is saved to output/models/{model_type}/actual_vs_predicted_test.png.


1.1 Exploratory Data Analysis

Design Decisions

Z columns are not missing data. Z variables encode the experimental recipe — set once before the run starts and recorded only at t=0. All subsequent rows carry NaN by design. extract_z_features filters to t=0 to extract the 13 scalar values cleanly.

Everything is aggregated to experiment level. The target is one scalar per experiment, but the input is a time series of varying length. Z variables are already scalars; W and X variables are aggregated using three statistics per variable: the last observed value (_last), the temporal mean (_mean), and the area under the curve (_auc, trapezoid rule). All three are computed in the EDA to enable direct comparison.

_last value features are excluded. Training experiments run for 7, 8, 9, 10, or 14 days; all test experiments are exactly 14 days. A last-value feature captures the measurement at the final time point, which occurs at different biological stages across experiments of different lengths — making the values incomparable between train and test. Mean and AUC aggregate over the full trajectory and are robust to this duration shift.

IVC is included as an engineered feature. IVC (Integral of Viable Cell Density) is the area under the X:VCD curve over time, computed with the trapezoid rule (unit: 10⁶ cells·day/mL). In mAb bioprocessing, antibody titer is classically proportional to the cumulative viable cell exposure: more cell-hours in the bioreactor means more antibody produced. IVC is a domain-standard feature that compresses the entire VCD time series into one interpretable scalar.

log(Y:Titer) is used for modelling. Raw titer has a right-skewed distribution (skewness ≈ +1.88). The log-transformed target is near-symmetric (skewness ≈ −0.055), which stabilises variance and benefits linear models. The transform is applied only at model training and evaluation — it is not stored in the data tables.

Key Insights

Each insight is traceable to a specific output of python src/eda.py.

Insight Console section Plot
log(Y:Titer) is near-Gaussian (skewness −0.055 vs +1.88 raw) Target Distribution (Y:Titer) eda_target_distribution.png
All 20 test experiments run exactly 14 days; train spans 7–14 Experiment Duration Distribution eda_duration_distribution.png
Z:ExpDuration matches the actual time-series length for all 120 experiments (normalization by planned duration is exact) Planned vs Actual Experiment Duration —
Z:ExpDuration is the strongest Z predictor (r ≈ +0.62) Z Variables — Pearson Correlation with Y:Titer eda_z_pearson.png
X:Lysed dominates aggregated feature correlations (mean r ≈ +0.69, AUC r ≈ +0.65) Aggregated W/X Features — Pearson Correlation with Y:Titer eda_aggregated_correlations.png
_last features are excluded due to duration shift — they capture measurements at different biological stages across experiments of different lengths Experiment Duration Distribution —
IVC correlates strongly with titer IVC vs Y:Titer eda_ivc_vs_titer.png
Q4 experiments separate from Q1 early in the run on X:VCD and X:Lysed — eda_x_timeseries.png

1.2 Feature Engineering

Design Decisions

The full feature matrix is built from three sources in src/features.py:

  • Z features (13) — scalars extracted from t=0 rows, one per experiment, representing the experimental recipe.
  • W/X aggregations (19) — mean and AUC per variable (4 W + 6 X = 10 variables × 2, minus X:VCD_auc). _last is excluded due to the train/test duration shift (see 1.1). X:VCD_auc is excluded because it is mathematically identical to IVC (both compute ∫ X:VCD dt via the trapezoid rule); keeping both would introduce a redundant feature with opposite-sign coefficients — a collinearity artefact. IVC is retained under its domain name.
  • IVC (1) — ∫ X:VCD dt per experiment (trapezoid rule).

AUC features are normalised by Z:ExpDuration (divided by the planned duration in days). Raw AUC grows proportionally with experiment length regardless of biology: a 14-day experiment accumulates roughly twice the AUC of a 7-day one even if the underlying dynamics are identical. This duration confounding is critical here because all test experiments run for exactly 14 days while the training set spans 7–14 days (durations: 7, 8, 9, 10, and 14 days; 90% shorter than 14 days). Dividing by duration converts AUC into a time-averaged rate, comparable across all experiment lengths.

Total: 33 features per experiment.


1.3 Model Training & Evaluation

Design Decisions

Model: regularised regression in log space. TiterPredictor wraps a StandardScaler + regressor pipeline. Training is performed on log(Y:Titer) (motivated by the near-Gaussian log-target from the EDA); predictions are back-transformed with exp() before scoring.

Why Ridge over ordinary least squares (OLS). With 100 experiments and 33 features, plain OLS tends to overfit and produce unstable coefficients. L2 regularisation (Ridge) reduces this by adding a penalty that shrinks coefficients toward zero, trading a small amount of bias for lower variance. The strength of the penalty is controlled by alpha: a larger alpha shrinks coefficients more aggressively (higher bias, lower variance); a smaller alpha approaches OLS (lower bias, higher variance). RidgeCV selects the optimal alpha from the explicit grid [0.01, 0.1, 1.0, 3.0, 5.0, 10.0, 30.0, 100.0, 1000.0] via 5-fold cross-validation. Selected alpha=3.0.

Why ElasticNet as a second model. ElasticNet combines L1 (sparsity) and L2 (stability) regularisation. The L1 term drives irrelevant feature coefficients exactly to zero, performing automatic feature selection. ElasticNetCV selects both alpha and l1_ratio via 5-fold CV over the extended grids [0.0001, 0.001, 0.005, 0.01, 0.05, 0.1, 0.5, 1.0, 5.0, 10.0] and [0.1, 0.5, 0.7, 0.9, 0.95, 1.0]. Selected alpha=0.005, l1_ratio=1.0 (pure LASSO) — zeros out 15 of 33 features and retains 18 informative predictors.

Why Random Forest as a third model. Random Forest is the natural non-parametric comparison point. It makes no linearity assumptions, captures feature interactions automatically, and provides Mean Decrease in Impurity (MDI) feature importances directly from the fitted trees. RandomForestRegressor(n_estimators=500, random_state=42) is wrapped in a GridSearchCV using the shared _INNER_CV splitter for consistency with Ridge and ElasticNet. The grid searches max_depth ∈ [None, 5, 10, 20], min_samples_leaf ∈ [1, 2, 5, 10], and max_features ∈ ["sqrt", 0.3, 0.5, 0.7, 1.0] — the hyperparameters most sensitive to overfitting on small datasets. With 100 training samples and 33 features, ensemble methods are typically disadvantaged relative to regularised linear models; the RF serves to confirm whether non-linear interactions provide additional signal beyond what linear models can capture. For this reason, RF is not used as a base for the adapted variant. Best hyperparameters selected by inner CV: {'max_depth': 10, 'max_features': 0.5, 'min_samples_leaf': 1}.

Adapted model (distribution shift mitigation). All test experiments run exactly 14 days, while only ~10% of the training set does. This creates a distribution shift with two consequences: Z:ExpDuration is constant on the test set (zero discriminative power there) and the model has seen few 14-day experiments across the full titer range. The adapted variant addresses this with two changes applied on top of the best linear base model (Ridge):

  1. Z:ExpDuration is removed from the input features. It is still used internally in build_feature_matrix to normalise AUC and IVC features, but not passed to the regressor.
  2. Inverse-frequency instance weighting: each experiment is weighted by 1 / freq(duration), normalised so the mean weight equals 1. This automatically upweights under-represented duration groups (14-day experiments) with no manually tuned multiplier.

The outer CV of the adapted model uses the same titer-stratified folds for a fair comparison, but its CV RMSE is not directly comparable to the base models: the validation folds contain mixed durations while the adapted model is optimised for 14-day experiments. A lower CV RMSE for a base model does not imply better performance on the all-14-day test set.

Evaluation: nested cross-validation (random_state=42, fully reproducible). Training and hyperparameter selection are kept strictly separate through two independent CV loops.

The outer loop (StratifiedKFold(5), stratified by titer quartile) splits the 100 training experiments into 5 folds of 80/20. Each held-out fold of 20 experiments is never seen during training or hyperparameter selection — it is used only to score the model. Averaging the 5 held-out scores gives an unbiased estimate of generalisation performance. Stratification by titer quartile ensures each fold contains a representative mix of low- and high-titer experiments, reducing fold-to-fold variance (plain KFold produced R² std ±0.299 due to titer clustering; stratification reduces this to ±0.068).

The inner loop operates on each fold's 80-experiment training set. RidgeCV tries every alpha in the explicit grid; ElasticNetCV searches over all (alpha, l1_ratio) combinations; RF uses GridSearchCV over depth, leaf size and feature fraction. All inner CV uses the shared _INNER_CV splitter (KFold(5, random_state=42)), so hyperparameter selection never touches the outer held-out set. The fold-level models are discarded after scoring.

After CV, a fresh TiterPredictor is trained on all 100 experiments for each model type. The inner CV selects the best hyperparameters on the full training set, and this final model is serialised to disk. The CV score estimates generalisation performance; the final model uses all available data for deployment. Metrics are computed in the original mg/L scale after back-transforming predictions:

  • RMSE / MAE (Root Mean Square Error / Mean Absolute Error) — absolute error in mg/L, directly interpretable as typical prediction error.
  • NRMSE — RMSE normalised by the full target range (283–4822 mg/L), expressed as a percentage. Provides scale-free context: an NRMSE of 7 % means the typical error is 7 % of the observable target range.
  • MAPE — mean absolute percentage error. Complements RMSE by expressing error relative to the actual titer of each experiment, useful when errors on low-titer runs are as important as errors on high-titer runs.

Diagnostic plots are saved in-sample per model to output/models/{model_type}/. Linear models (ridge, elasticnet, ridge_adapted) produce coefficients.png; the RF model produces feature_importance.png (MDI).

Plot Description
actual_vs_predicted.png Scatter of predicted vs actual titer (mg/L) with perfect-prediction diagonal
residuals.png Residuals (ŷ − y) vs actual titer in mg/L — checks for systematic bias
log_residuals.png Residuals in log space (log ŷ − log y) vs log y — checks bias in the modelling scale
coefficients.png Standardised coefficients (log mg/L per std dev), sorted by absolute value (linear models)
feature_importance.png MDI feature importances sorted by value (RF only)

Key Results (train set, 5-fold stratified CV)

Metric Ridge ElasticNet RF Ridge_Adapted†
RMSE 276.0 ± 95.3 mg/L 298.0 ± 117.8 mg/L 400.4 ± 105.3 mg/L 308.5 ± 122.7 mg/L
NRMSE 6.1 ± 2.1 % 6.6 ± 2.6 % 8.8 ± 2.3 % 6.8 ± 2.7 %
MAE 188.5 ± 33.5 mg/L 197.0 ± 49.0 mg/L 242.5 ± 44.1 mg/L 211.5 ± 41.5 mg/L
MAPE 14.9 ± 2.1 % 14.8 ± 2.4 % 17.9 ± 3.4 % 16.7 ± 1.2 %
R² 0.854 ± 0.068 0.829 ± 0.082 0.709 ± 0.086 0.822 ± 0.094

The RF result confirms that no non-linear interactions above what regularised linear models capture are detectable with 100 training samples. With more data, RF or gradient boosting would be the natural next step.

† CV RMSE for the adapted model is not directly comparable — see Design Decisions above.

Ridge is selected as the final model (lowest mean CV RMSE). Selected alpha=3.0. ElasticNet selected alpha=0.005, l1_ratio=1.0 (pure LASSO — 15 of 33 features zeroed out, 18 retained).

Coefficient interpretation. Model coefficients are in log(mg/L) per standard deviation of each feature: a coefficient c means that a one-std-dev increase in the feature multiplies the predicted titer by exp(c). Ridge retains all 33 features with non-zero but shrunken coefficients; ElasticNet (pure LASSO, l1_ratio=1.0) explicitly zeroes 15 features and selects 18. The top-5 by magnitude for each model:

Feature Ridge coef ElasticNet coef Direction
X:Lac_mean +0.216 +0.406 positive — strongest signal in both
X:Lac_auc +0.183 ≈ 0 (zeroed) collinear with X:Lac_mean; L1 drops it
Z:ExpDuration +0.169 +0.222 positive — longer runs produce more
X:Gln_mean (glutamine) −0.135 −0.235 negative — consistent across both models
W:FeedGln_auc +0.101 +0.136 positive — consistent across both models
X:Glc_mean (glucose) — −0.201 negative — not in Ridge top 5; ElasticNet explicitly selects it

The directionality is fully consistent across both models. X:Lac_mean is the dominant predictor in both: higher mean lactate reflects high metabolic activity and antibody productivity. X:Lac_auc is dropped by ElasticNet (L1 selects the more informative mean statistic when the two are collinear). X:Gln_mean and X:Glc_mean are negative — high residual substrate concentrations suggest suboptimal consumption and lower productivity. W:FeedGln_auc is positive in both, reflecting that sustained glutamine feeding extends productive culture.


Part 2: Inference Server

A FastAPI microservice that serves the trained model via two endpoints defined in inference_server_spec.yml.

Setup

Install server dependencies (requires a trained model from Part 1):

pip install -r requirements-server.txt

requirements-server.txt is the unified runtime requirements file: it contains the ML pipeline packages needed at inference time (pandas, numpy, scipy, scikit-learn, joblib, pyyaml) together with the server framework dependencies (fastapi, uvicorn, pydantic). requirements.txt retains the full development dependencies (formatters, linters, visualisation libraries) and is used by the devcontainer postCreateCommand.

Running

Two options are available to start the server (reads model_path from config.yaml). Option A is recommended for development; Option B runs the exact production image.

Option A — local development

PYTHONPATH=src uvicorn server.app:app --host 0.0.0.0 --port 8080 --reload

Option B — Docker

docker build -t datahow-server .
docker run -p 8080:8080 datahow-server

Both options expose the same two endpoints on port 8080:

  • GET /health — returns {"status": "ok"}
  • POST /predict — accepts a JSON payload with timestamps and values, returns predicted titer plus model metadata (model_type, model_version)

Tests (no running server required):

python -m pytest tests/ -v

Endpoints

GET /health

Liveness check. Returns {"status": "ok"} with HTTP 200. No authentication or model state is required.

POST /predict

Accepts a JSON payload describing one bioprocess experiment and returns the predicted final titer.

A full example request is provided in examples/predict_request.json (14-day experiment from the OpenAPI spec, 15 time points spanning days 0–14). Test the running server with:

curl -X POST http://localhost:8080/predict \
  -H "Content-Type: application/json" \
  -d @examples/predict_request.json

Expected response (examples/predict_response.json):

{
  "titer_predicted_mg_L": 1993.8552,
  "model_type": "ridge",
  "model_version": "465a92ceee45"
}

Request schema

  • timestamps: strictly increasing array of time points in days (non-empty, length n). Non-consecutive spacing is supported — AUC features use the trapezoid rule over the actual time values.
  • values: variable arrays keyed by prefix convention:
    • Z: — scalar recipe parameters, single-element arrays
    • W: and X: — time-series arrays of length n

All 23 variables used during training are required (Z_COLS, W_COLS, X_COLS from data_loading.py). The implementation satisfies all spec-mandated validation rules and adds further constraints to prevent silent feature corruption. Invalid payloads are rejected with HTTP 422 (FastAPI/Pydantic standard for schema and content validation failures; the OpenAPI spec names 400, but 422 is semantically correct for well-formed requests that fail content validation).

Rules required by the spec:

  • all required Z, W, and X variables must be present
  • Z variables must be single-element arrays
  • W and X arrays must have length equal to len(timestamps)

Additional constraints enforced by the implementation:

  • unknown top-level fields are rejected (extra="forbid")
  • timestamps must be non-empty and strictly increasing (no duplicates, no out-of-order values)
  • Z:ExpDuration must be greater than 0
  • Z:ExpDuration must equal max(timestamps) — AUC features are normalised by Z:ExpDuration in build_feature_matrix; a mismatch would corrupt features silently at inference time

Response schema

  • titer_predicted_mg_L: predicted final product titer in mg/L
  • model_type: model variant used ("ridge", "ridge_adapted", "elasticnet", "rf")
  • model_version: SHA-256 fingerprint (12 hex chars) of the model.joblib file loaded at startup — ties each prediction to the exact serialised artefact

Architecture

POST /predict
    ↓
PredictRequest (Pydantic v2 validation)
    ↓
inference.predict_from_payload()
    ├── _payload_to_raw_df()   — reshape JSON → raw DataFrame (one row per time point)
    ├── build_feature_matrix() — aggregate to experiment-level feature vector (33 features)
    └── TiterPredictor.predict() — StandardScaler → Ridge → exp() → mg/L
    ↓
PredictResponse

The payload is converted to the same raw DataFrame format used during training, so the entire feature engineering pipeline (build_feature_matrix) runs unchanged. The exclude_duration flag is read from the serialised model, so the correct feature set is used automatically regardless of which model variant is configured in config.yaml.

The predictor is loaded once at startup via FastAPI's lifespan context manager and injected into each request via a Depends dependency, enabling clean override in tests without touching the lifespan.

Serialisation

Deserialization (JSON → Python) and serialization (Python → JSON) are handled transparently by Pydantic v2 and FastAPI:

  • Incoming request: FastAPI passes the JSON body to Pydantic, which coerces JSON numbers to float, arrays to list[float], and objects to dict[str, list[float]]. Unknown top-level fields are rejected with HTTP 422 (extra="forbid"). The parsed model is immutable (frozen=True) — no code path can mutate it after construction.
  • Outgoing response: FastAPI calls PredictResponse.model_dump() and serialises the result to JSON. PredictResponse is immutable (frozen=True). A math.isfinite guard in predict_from_payload rejects non-finite prediction values before the response is constructed, so they are caught by the existing post_predict exception handler and returned as a clean HTTP 500.

Testing

29 integration tests cover both endpoints, all validation paths, and internal error handling. The test suite does not require a trained model file: a synthetic TiterPredictor is fitted in conftest.py and injected via app.dependency_overrides. The synthetic training set consists of 20 copies of the spec example experiment's feature vector with small Gaussian noise (σ=0.01) added to each of the 33 features, paired with random uniform targets in [500, 2000] — enough to produce a well-defined Ridge fit with the correct feature schema, without requiring the real dataset.

The tests cover only Part 2 (inference server). Unit tests for the Part 1 pipeline (features.py, model.py) were not implemented within the challenge time frame and are listed under Possible Extensions.

tests/test_server.py::TestHealth::test_returns_200
tests/test_server.py::TestHealth::test_body_contains_ok_status
tests/test_server.py::TestPredict::test_returns_200
tests/test_server.py::TestPredict::test_response_contains_titer_key
tests/test_server.py::TestPredict::test_prediction_is_positive_float
tests/test_server.py::TestPredict::test_prediction_is_deterministic
tests/test_server.py::TestPredict::test_response_contains_model_type
tests/test_server.py::TestPredict::test_model_type_is_non_empty_string
tests/test_server.py::TestPredict::test_model_type_matches_fixture
tests/test_server.py::TestPredict::test_response_contains_model_version
tests/test_server.py::TestPredict::test_model_version_is_non_empty_string
tests/test_server.py::TestPredict::test_model_version_matches_fixture
tests/test_server.py::TestPredictValidation::test_missing_z_variable_returns_422
tests/test_server.py::TestPredictValidation::test_missing_x_variable_returns_422
tests/test_server.py::TestPredictValidation::test_missing_w_variable_returns_422
tests/test_server.py::TestPredictValidation::test_empty_timestamps_returns_422
tests/test_server.py::TestPredictValidation::test_z_variable_multi_element_array_returns_422
tests/test_server.py::TestPredictValidation::test_x_variable_length_mismatch_returns_422
tests/test_server.py::TestPredictValidation::test_missing_timestamps_field_returns_422
tests/test_server.py::TestPredictValidation::test_missing_values_field_returns_422
tests/test_server.py::TestPredictValidation::test_empty_body_returns_422
tests/test_server.py::TestPredictValidation::test_non_increasing_timestamps_returns_422
tests/test_server.py::TestPredictValidation::test_duplicate_timestamps_returns_422
tests/test_server.py::TestPredictValidation::test_zero_exp_duration_returns_422
tests/test_server.py::TestPredictValidation::test_negative_exp_duration_returns_422
tests/test_server.py::TestPredictValidation::test_exp_duration_mismatch_returns_422
tests/test_server.py::TestPredictValidation::test_extra_top_level_field_returns_422
tests/test_server.py::TestPredictErrorPaths::test_prediction_failure_returns_500
tests/test_server.py::TestPredictErrorPaths::test_non_finite_prediction_returns_500

Possible Extensions

Unit tests for the Part 1 pipeline. The current test suite covers the inference server only. The highest-value additions would be: (1) unit tests for build_feature_matrix — verify feature count (33), column names, and AUC normalisation by Z:ExpDuration; (2) unit tests for TiterPredictor — verify the save/load roundtrip and that predict returns positive values after back-transforming from log space.

Finer hyperparameter tuning. The current grids use explicit value arrays chosen to cover the plausible range without excessive computation. A natural next step is automated grid densification: RidgeCV accepts arbitrarily many alphas at no extra cost (SVD — Singular Value Decomposition — based), so np.logspace(-3, 4, 100) would find near-continuous optima. For ElasticNet and RF, RandomizedSearchCV over continuous distributions (e.g. loguniform for alpha, randint for tree depth) would explore a much larger space within a fixed compute budget.

MLP on engineered features. The simplest non-linear extension is a shallow MLP (Multi-Layer Perceptron) trained directly on the 33 features. ElasticNet's feature selection (18 non-zero out of 33) could serve as a preprocessing step, reducing dimensionality before the MLP to limit overfitting. With 100 training samples overfitting risk is high; sklearn's MLPRegressor offers only L2 regularisation, while a PyTorch implementation would enable dropout and early stopping, making this approach more viable even on small datasets.

1D CNN (Convolutional Neural Network) on raw time series. A 1D convolutional network could process the W and X time series directly, learning temporal features without manual aggregation. It would require a hybrid architecture (CNN branch for W/X sequences + dense branch for Z scalars), padding or masking to handle variable-length inputs (7–14 days), and PyTorch as an additional dependency. With 100 samples the risk of overfitting is substantial; this extension becomes practical with a larger dataset of bioprocess experiments.


AI Tools & Development Process

Approach

I acted as architect and domain decision-maker throughout. All structural choices — what to build, in what order, and what to include or exclude — originated from me. Claude Code's role was to implement those decisions, surface technical constraints (API changes, type errors, style violations), and propose alternatives when asked — never to set direction autonomously.

Development Phases and Prompts

Phase 1 — Problem understanding

"Analyze the challenge PDF and the dataset. Explain the problem, the variable groups, and what biological meaning each prefix (Z/W/X/Y) carries."

I built a shared understanding of the domain before committing to any design choice. This grounded all downstream decisions in the actual data structure and biological context. The Variable Interaction Overview Mermaid diagram was proposed and produced in this phase.

Phase 2 — EDA pipeline

"Start with an EDA. I want: target distribution, Pearson correlation of each variable with titer, and time-series profiles of W and X grouped by titer quartile."

I defined the scope of the analysis: correlation plots, quartile-stratified time-series profiles, histogram and Q-Q plots for the target.

Phase 3 — Aggregation strategy

"Aggregate everything to experiment level. For Z variables use only the t=0 rows — those are the recipe parameters. For W and X, reduce the time series to scalars using mean and AUC. Skip the last-value — it is not comparable across experiments of different durations."

I identified the train/test duration-shift problem and decided the aggregation scheme before implementation started.

Phase 4 — Domain knowledge and feature engineering

"What domain-specific features or tricks are standard for this type of bioprocess data? What does the literature typically use to predict antibody titer from cell culture time series?"

This led to the introduction of IVC (Integral of Viable Cell Density) — the trapezoid integral of X:VCD over time — as a domain-standard engineered feature that captures cumulative cell exposure and is classically proportional to final antibody titer.

Phase 5 — Model selection and evaluation framework

"Add ElasticNet as a second model type. Design a nested CV framework: outer StratifiedKFold(5) on titer quartile for generalisation estimation, inner 5-fold CV for hyperparameter selection (RidgeCV, ElasticNetCV). After CV, train a separate final model on all 100 experiments. Print a side-by-side comparison and automatically update config.yaml to point to the model with the lower RMSE."

Deciding which models to compare first, then designing the evaluation framework around them, keeps the CV design tightly scoped. The nested structure separates generalisation estimation (outer) from hyperparameter selection (inner): inner CV operates only on each fold's training set and never touches the held-out data. The saved model is trained from scratch on all 100 experiments.

Phase 6 — Test-set inference and configuration

"Write a predict.py script for test-set inference. Read all paths from a config.yaml. Save predictions with Y:Titer_actual and Y:Titer_predicted columns side by side. Print evaluation metrics to stdout only — not in the CSV."

This produced predict.py + config.yaml as the inference entry point. Updating targets_path in config.yaml to the received test targets file and rerunning is the only step needed to evaluate against real test targets.

Phase 7 — Distribution shift and adapted model

"The test set is entirely made of 14-day experiments while only ~10% of the training set does. Implement an adapted model variant: exclude Z:ExpDuration from the input features and add inverse-frequency instance weighting to upweight 14-day experiments."

The EDA confirmed that Z:ExpDuration is the strongest Z predictor in the training set (r ≈ +0.62), but it is constant on the test set — providing no discriminative signal there. The adapted variant addresses this with two changes on top of the best linear base model: (1) Z:ExpDuration is removed from the input features (still used internally for AUC normalisation); (2) inverse-frequency instance weighting upweights under-represented duration groups automatically. The exclude_duration flag is serialised inside the model so that inference code reconstructs the correct feature matrix without additional configuration.

Phase 8 — Random Forest and hyperparameter grid refinement

"Add Random Forest as a third base model using GridSearchCV with the shared inner CV splitter. RF should not be used as a base for the adapted model. Add MDI feature importance plots. Also extend the hyperparameter grids for Ridge and ElasticNet."

Refining the grids changed the winner: Ridge improved from 303.3 to 276.0 mg/L by selecting alpha=3.0 — a value not reachable with the original 6-point grid. ElasticNet shifted to a smaller alpha (0.005) with pure LASSO (l1_ratio=1.0) — more aggressive feature selection but higher CV RMSE than Ridge. RF underperforms on this 100-sample dataset (RMSE 400.4 mg/L) even after tuning max_depth, confirming that linear models are more appropriate at this data scale.

Phase 9 — Response schema design

"Design the POST /predict response schema. Keep the flat columnar request format that mirrors the CSV Z:/W:/X: prefix convention. Enrich the response with model_type and model_version: model_type derived from predictor.model_type and the exclude_duration flag; model_version as the first 12 hex chars of the SHA-256 of model.joblib."

model_type emits "ridge_adapted" when exclude_duration=True. model_version is the first 12 hex chars of the SHA-256 hash of the serialised model.joblib file, tying each prediction to the exact trained artefact.

Phase 10 — Example payload files

"Add an example input payload and an example response as local files in an examples/ directory."

examples/predict_request.json uses the spec example experiment (15 timestamps, all 23 variables from inference_server_spec.yml). examples/predict_response.json records the real Ridge output (1993.8552 mg/L, version 465a92ceee45) so that the expected response is traceable to the exact trained artefact.

Phase 11 — Additional input validations

"Add additional input validations to PredictRequest: enforce strictly increasing timestamps, Z:ExpDuration > 0, and Z:ExpDuration == max(timestamps) to prevent silent feature corruption at inference time."

The decision was to enforce Z:ExpDuration == max(timestamps) at the API boundary rather than changing the normalisation denominator in build_feature_matrix (which would require retraining). Three validators were added to PredictRequest in src/server/schemas.py: (1) timestamps strictly increasing, checked in the timestamps_non_empty field validator; (2) Z:ExpDuration > 0; (3) Z:ExpDuration == max(timestamps) within 1e-6 — both in validate_values. Silently wrong predictions are replaced by explicit 422 responses. This phase produced the first 26 integration tests.

Phase 12 — DTO (Data Transfer Object) serialisation hardening

"Harden the Pydantic DTOs: add extra='forbid' and frozen=True to PredictRequest, frozen=True to PredictResponse, and a math.isfinite guard in predict_from_payload to catch non-finite prediction values before the response is constructed."

model_config = ConfigDict(extra="forbid", frozen=True) added to PredictRequest: unknown top-level fields are rejected with HTTP 422 and the parsed model is immutable after construction. model_config = ConfigDict(frozen=True) added to PredictResponse. A math.isfinite guard in predict_from_payload catches non-finite prediction values before the response is constructed and returns a clean HTTP 500 via the existing exception handler. One test added for extra="forbid" brings the total to 27.

Phase 13 — Dockerization improvements

"Containerise the inference server: python:3.13-slim base image, a unified requirements-server.txt that includes both ML runtime and FastAPI/uvicorn/pydantic dependencies, and a HEALTHCHECK using urllib.request."

requirements-server.txt was rewritten as the unified runtime requirements file, merging ML packages with the server framework and removing dev-only dependencies from the container. devcontainer.json postCreateCommand updated to install both requirements.txt and requirements-server.txt explicitly.

About

A machine learning pipeline for predicting monoclonal antibody titer from bioprocess data, including feature engineering, model evaluation, and a FastAPI inference service.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages