Skip to content
#

model-benchmarking

Here are 32 public repositories matching this topic...

A modular deep learning evaluation framework for benchmarking multiple CNN architectures across varied optimization strategies and training configurations. Built for scalable experimentation and transferability to real-world image classification tasks.

  • Updated Jun 19, 2025
  • Jupyter Notebook

A reproducible, leak-free machine learning benchmarking lab for regression model comparison, cross-validation, diagnostics, experiment tracking, and interactive Streamlit-based inference using the California Housing dataset.

  • Updated Oct 2, 2026
  • Python
LLM-as-Judge

A Streamlit web app that uses a Groq-powered LLM (Llama 3) to act as an impartial judge for evaluating and comparing two model outputs. Supports custom criteria, presets like creativity and brand tone, and returns structured scores, explanations, and a winner. Built end-to-end with Python, Groq API, and Streamlit.

  • Updated Jul 4, 2026
  • HTML

Causal analysis framework using Double Machine Learning to quantitatively isolate the effect of model size on deep learning performance while controlling for confounders such as dataset size, training time, and hyperparameters.

  • Updated Feb 14, 2026
  • Python

Add this topic to your repo

To associate your repository with the model-benchmarking topic, visit your repo's landing page and select "manage topics."

Learn more