Code, data and experiment configurations for encoder-only and decoder-only large language models applied to multi-label emotion classification of mobile app reviews, including data-imbalance mitigation and synthetic data augmentation.
The human-labelled ground truth and annotation guidelines are those introduced in Motger et al. (2025).
The best encoder-only configurations found by this project's experiments, refit on the full training pool (all cross-validation folds combined) plus the same generative-augmentation synthetic pool used in the imbalance-mitigation sweep, are published on the Hugging Face Hub:
| Model | Formulation | Macro-F1 (10-fold CV) | Link |
|---|---|---|---|
emotion-roberta-large-multilabel-genai-bce |
Multi-label (shared head), RoBERTa-large, generative augmentation ( |
0.591 ± 0.054 | huggingface.co/quim-motger/emotion-roberta-large-multilabel-genai-bce |
emotion-bert-base-binary-ensemble |
Binary ensemble (9 independent classifiers), BERT-base, generative augmentation ( |
0.530 ± 0.074 | huggingface.co/quim-motger/emotion-bert-base-binary-ensemble |
The binary-ensemble repo holds nine independent single-label checkpoints, one subfolder per
emotion (Joy/, Trust/, ... Neutral/) — load each with
AutoModelForSequenceClassification.from_pretrained(repo_id, subfolder="Joy"). See each
model's card for the full decision rule, usage example, and training recipe.
Both were produced by scripts/train_final_model.py and
scripts/train_final_binary_ensemble.py, which refit the winning imbalance-mitigation configuration on the
full training pool and push directly to the Hub (--push --repo-id ...); re-run either to
reproduce or update the published checkpoints.
- Python 3.10+
- Ollama for local open-source generative models
- API keys for proprietary models (optional, depending on the experiment)
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -U pip
pip install -e .
cp .env.example .env # then set keys as neededEnvironment variables are documented in .env.example.
| Path | Description |
|---|---|
configs/ |
Experiment configuration (models, training, imbalance) |
src/emotion_cls/ |
Source package and emotion-cls CLI |
Datasets/ |
Ground-truth CSV, annotation guidelines, and synthetic review corpora |
scripts/ |
Auxiliary data-preparation scripts |
human_eval/ |
Human-as-judge validation materials (annotation files, not experiment code) |
outputs/ |
Experiment outputs (created at runtime) |
models/ |
Fine-tuned checkpoints (created at runtime; gitignored) |
Encoder-only and decoder-only models are declared in configs/models.yaml.
emotion-cls list-modelsEncoder fine-tuning uses Hugging Face Transformers (PyTorch). Open-source decoder-only models are served with Ollama; proprietary models use their official APIs.
Default settings (10-fold multilabel stratified CV, threshold-based label assignment capped at 3, emotion set, hyperparameters) are in configs/default.yaml. Override via CLI flags or by editing the YAML files.
# Multi-label fine-tuning
emotion-cls train-encoder --encoder bert-base-cased --head multilabel
emotion-cls train-encoder --encoder roberta-large --head multilabel
# Binary ensemble (one classifier per emotion; threshold assembly, capped at 3)
emotion-cls train-encoder --encoder bert-base-cased --head binaryResults are written under outputs/encoder_multilabel/ and outputs/encoder_binary/.
Prompting strategies: zero_shot, few_shot_guidelines, few_shot_guidelines_dataset.
Prompts are built from Datasets/guidelines/Annotation Guidelines.txt (the official annotation guidelines).
emotion-cls classify-decoder --decoder gemma3-4b --strategy zero_shot
emotion-cls classify-decoder --decoder gemma3-4b --strategy zero_shot --temperature 0.3
emotion-cls classify-decoder --decoder claude-haiku-4-5 --strategy few_shot_guidelines_datasetTemperature is a tunable decoding factor (--temperature; default 0.0 from configs/default.yaml).
Grid candidates are listed under decoding.temperature_grid (0.0, 0.3, 0.7).
Each temperature writes to its own resumable directory
outputs/decoder_classify/<decoder>/<strategy>/t<temp>/.
--imbalance |
Description |
|---|---|
none |
Baseline multilabel BCE |
bce_pos_weight |
pos_weight[c] = neg_c / pos_c |
bce_weight |
w[c] = N / (|E| · pos_c) |
focal |
Focal loss with class-wise α |
adaptive_focal |
Adaptive Focal Loss |
undersample |
Cap positives per emotion (--undersample-cutoff) |
genai_aug |
Inject synthetic reviews (--aug-inject-n, --synthetic-ml-path) |
Methods may be combined with commas (e.g. undersample,bce_pos_weight).
emotion-cls train-encoder --encoder bert-base-cased --imbalance bce_pos_weight
emotion-cls train-encoder --encoder bert-base-cased --imbalance undersample --undersample-cutoff 150
emotion-cls sweep-undersample --encoder bert-base-cased --cutoffs 50,100,150,200,250,300,350
emotion-cls train-encoder --encoder bert-base-cased \
--imbalance genai_aug,bce_pos_weight \
--aug-inject-n 100 \
--synthetic-ml-path Datasets/synthetic_multilabel.csvBuild a multilabel synthetic CSV from provider folders under Datasets/:
python scripts/build_synthetic_multilabel.py \
--strategy few_shot_guidelines_dataset \
--genai Claude \
--n-per-emotion 100 \
--out Datasets/synthetic_multilabel.csv# Rank existing synthetic corpora (diversity / novelty / on-emotion + Borda)
emotion-cls rank-augmentation
# Generate until majority-class parity
emotion-cls generate --decoder claude --emotion Fear --strategy few_shot_guidelines_dataset
emotion-cls generate-parity --decoder claude --strategy few_shot_guidelines_datasetSynthetic outputs are stored under Datasets/generated/ and are not mixed into Datasets/GroundTruth.csv.
emotion-cls predict \
--model-path models/<checkpoint> \
--input reviews.txt \
--output predicted_reviews.csvconfigs/default.yaml— data paths, CV, training hyperparametersconfigs/models.yaml— encoder and decoder cataloguesconfigs/imbalance.yaml— imbalance method definitions
emotion-cls export-run-config --out outputs/resolved_config.yamlRuns are resumable by default (experiment.resume: true in configs/default.yaml).
Re-executing the same command skips finished work and continues from the last checkpoint.
| Experiment | Resume unit | Periodic flush |
|---|---|---|
| Encoder multilabel | completed fold_N/metrics.json |
after each fold (+ HF epoch checkpoints) |
| Encoder binary | fold metrics; within fold, probs.npy / emotions_done.json per emotion |
after each emotion |
| Decoder classify | fold metrics; within fold, per-sentence predictions.csv |
after every sentence |
| Generation | existing synthetic CSV row count | after every batch |
| Undersample sweep | per-cutoff directory undersample_c<N>/ |
same as encoder |
Each run directory also stores:
run_meta.json— config snapshot, start/end timestamps, run idevents.jsonl— append-only event log (fold start/done, LLM calls, errors)progress.json— completed folds/unitsusage_totals.json— aggregated prompt/completion tokens and latency (API / Ollama)
Force a clean re-run with --no-resume.
emotion-cls train-encoder --encoder bert-base-cased --head multilabel
emotion-cls classify-decoder --decoder gemma3-4b --strategy zero_shot
# after a crash, the same commands continue where they left offPlease cite the ground-truth dataset paper when using this package:
Q. Motger, M. Oriol, M. Tiessler, X. Franch and J. Marco, "What About Emotions? Guiding Fine-Grained Emotion Extraction from Mobile App Reviews," 2025 IEEE 33rd International Requirements Engineering Conference (RE), Valencia, Spain, 2025, pp. 6-18, doi: 10.1109/RE63999.2025.00012.