Software
Open-source simulators, hardware designs, benchmarks, and analysis tools developed by our group. Most are available on GitHub at github.com/harvard-acc.
Jump to: Featured · Accelerator Simulation & Design · Hardware Accelerator Designs · Memory, Chiplet & Power Modeling · ML Systems & Recommendation · Reliability & Resilience · Benchmarks & Workloads · Workload Characterization & Analysis · Compilers & CPU Simulation
Featured
Aladdin
Aladdin is our pre-RTL power and performance simulator for hardware accelerators.
gem5-Aladdin
gem5-Aladdin is an integration of the Aladdin accelerator simulator with the gem5 system simulator to enable simulation of end-to-end accelerated workloads on SoCs.
GoldenEye
GoldenEye is a functional simulator with fault injection capabilities for common and emerging numerical formats, implemented for the PyTorch deep learning framework.
Smaug
SMAUG is a deep learning framework that enables end-to-end simulation of DL models on custom SoCs with a variety of hardware accelerators.
DreamRAM
DreamRAM is a configurable modeling and design space exploration tool for custom 3D die-stacked DRAM architectures based on HBM.
NVMExplorer
NVMExplorer is a cross-stack design space exploration framework for evaluating and comparing on-chip memory solutions including emerging, embedded non-volatile memories.
Accelerator Simulation & Design
Aladdin
Aladdin is our pre-RTL power and performance simulator for hardware accelerators.
gem5-Aladdin
gem5-Aladdin is an integration of the Aladdin accelerator simulator with the gem5 system simulator to enable simulation of end-to-end accelerated workloads on SoCs.
Smaug
SMAUG is a deep learning framework that enables end-to-end simulation of DL models on custom SoCs with a variety of hardware accelerators.
Trireme
Trireme is a framework for automatically exploring hierarchical multi-level parallelism for domain-specific hardware acceleration.
Hardware Accelerator Designs
EdgeBERT
EdgeBERT is a HW/SW co-design enabling sentence-level energy optimizations for latency-aware multi-task NLP inference.
ESP Systolic Array Accelerator
ESP Systolic Array Accelerator is an 8-bit integer, weight-stationary systolic array accelerator designed with the Catapult HLS flow in SystemC, with a reconfigurable array size (8, 16, or 32) and an AXI configuration interface.
FlexASR
FlexASR is an AXI-programmable hardware accelerator for attention-based seq-to-seq networks. It can be configured to accelerate end-to-end RNN, GRU or LSTM models with attention mechanisms (e.g. Listen-Attend-and-Spell models).
Memory, Chiplet & Power Modeling
CASCADE
CASCADE (Composable Analytical System of Chiplets for AI Devices at the Edge) is a design space exploration framework for heterogeneous chiplet-based System-in-Package architectures for edge ML.
DreamRAM
DreamRAM is a configurable modeling and design space exploration tool for custom 3D die-stacked DRAM architectures based on HBM.
McPAT CPU Models
We have developed McPAT power models for a recent high-performance multicore CPU.
NVMExplorer
NVMExplorer is a cross-stack design space exploration framework for evaluating and comparing on-chip memory solutions including emerging, embedded non-volatile memories.
ML Systems & Recommendation
DeepRecSys
DeepRecSys provides an end-to-end infrastructure to study and optimize at-scale neural recommendation inference.
RecPipe
RecPipe provides an end-to-end system to study and jointly optimize recommendation models and hardware for at-scale inference.
Reliability & Resilience
Benchmarks & Workloads
BayesSuite
BayesSuite is a collection of Bayesian inference workloads written in Stan framework.
Fathom
Fathom is a collection of workloads for benchmarking modern machine learning techniques.
MachSuite
MachSuite is a benchmark suite for high-level synthesis and accelerator-centric architectures.
Workload Characterization & Analysis
CHAMPVis
CHAMPVis, Comparative Hierarchical Analysis of Microarchitectural Performance Visualization.
LLVM-Tracer
LLVM-Tracer is an LLVM instrumentation pass to print out a dynamic LLVM IR execution trace, including dynamic values and memory addresses.
WIICA
WIICA is the Workload ISA Independent Characterization for Applications tool.