"Optimizing Performance and Energy Efficiency in Massively Parallel Systems" PhD Dissertation repository.
-
Updated
Oct 14, 2022
"Optimizing Performance and Energy Efficiency in Massively Parallel Systems" PhD Dissertation repository.
Immix mark-region garbage collector utilizing hierarchical block and line marks, fine-grained object allocation, and opportunistic mark-sweep-evacuate defragmentation.
Generational GC with nursery and mature generations using write-barrier card table tracking
jemalloc-style high concurrency slab allocator with multi-size classes, bitmap run tracking, arena partitioning, and dirty page decay management.
Tri-color abstraction mark-and-sweep garbage collection engine with worklist traversal
jemalloc-style high concurrency slab allocator with multi-size classes, bitmap run tracking, arena partitioning, and dirty page decay management.
Generational GC with nursery and mature generations using write-barrier card table tracking
Reference counting engine with Bacon-Rajan trial deletion cycle collection
Tri-color abstraction mark-and-sweep garbage collection engine with worklist traversal
Reference counting engine with Bacon-Rajan trial deletion cycle collection
Immix mark-region garbage collector utilizing hierarchical block and line marks, fine-grained object allocation, and opportunistic mark-sweep-evacuate defragmentation.
XiTAO is a lightweight layer built on top of modern C++ features with the goals of being low-overhead and serving as a development platform for testing scheduling and resource management algorithms.
Execution Constraint Engine (ECE) is a runtime decision layer for multi-step LLM workflows. ECE (v1) focuses on cost constraints, acting as a guardrail for unbounded execution in loops, agents, and multi-step systems. It evaluates projected cost before each step and stops execution before exceeding a defined budget.
Runtime-core research for long-lived AI surfaces: worker ownership, transaction scheduling, and bounded projection.
Relational Time Engine (RTE): runtime density regulation for compute-efficient transformer inference. Demonstrates up to 75% layer reduction with improved latency and throughput.
Sessions, transactions, and version history for your knowledge graph. Operate without becoming the authority.
Exploratory AI infrastructure project modeling semantic KV-cache orchestration, memory tiering, and HBM/CXL movement tradeoffs for long-context LLM inference.
A systems research platform for semantic KV-cache orchestration, topology-aware memory placement, distributed prefix reuse, and rack-scale inference memory simulation.
Deterministic intermediate representation for AI agents — compile, verify, execute, and replay structured intent.
Predictive SRAM–DRAM–SSD memory fabric for low-cost AI inference, long-context KV cache tiering, MoE expert staging, and software-defined memory orchestration.
To associate your repository with the runtime-systems topic, visit your repo's landing page and select "manage topics."