Software

Open-source simulators, hardware designs, benchmarks, and analysis tools developed by our group. Most are available on GitHub at github.com/harvard-acc.

Jump to: Featured · Accelerator Simulation & Design · Hardware Accelerator Designs · Memory, Chiplet & Power Modeling · ML Systems & Recommendation · Reliability & Resilience · Benchmarks & Workloads · Workload Characterization & Analysis · Compilers & CPU Simulation

Accelerator Simulation & Design

Aladdin

Aladdin is our pre-RTL power and performance simulator for hardware accelerators.

Future heterogeneous architecture with CPU cores, a GPGPU, shared cache and a sea of accelerators (from the Aladdin paper)

gem5-Aladdin

gem5-Aladdin is an integration of the Aladdin accelerator simulator with the gem5 system simulator to enable simulation of end-to-end accelerated workloads on SoCs.

Example SoC modeled in gem5-Aladdin: CPUs, caches, DMA engine, and cache- and scratchpad-based accelerators

Smaug

SMAUG is a deep learning framework that enables end-to-end simulation of DL models on custom SoCs with a variety of hardware accelerators.

Overview of SMAUG execution flow, from the Python front end through tiling and scheduling to accelerator and CPU back ends

Trireme

Trireme is a framework for automatically exploring hierarchical multi-level parallelism for domain-specific hardware acceleration.

Overview of the Trireme methodology for selecting hardware accelerators and multi-level parallelism

Hardware Accelerator Designs

EdgeBERT

EdgeBERT is a HW/SW co-design enabling sentence-level energy optimizations for latency-aware multi-task NLP inference.

EdgeBERT overview: early-exit transformer encoder with network pruning, floating-point quantization and adaptive attention span

ESP Systolic Array Accelerator

ESP Systolic Array Accelerator is an 8-bit integer, weight-stationary systolic array accelerator designed with the Catapult HLS flow in SystemC, with a reconfigurable array size (8, 16, or 32) and an AXI configuration interface.

Schematic of the HLS weight-stationary systolic array accelerator with weight, input and output buffers and an AXI interface

FlexASR

FlexASR is an AXI-programmable hardware accelerator for attention-based seq-to-seq networks. It can be configured to accelerate end-to-end RNN, GRU or LSTM models with attention mechanisms (e.g. Listen-Attend-and-Spell models).

FlexASR accelerator architecture with processing elements, global buffer, activation unit and AXI bus

Memory, Chiplet & Power Modeling

CASCADE

CASCADE (Composable Analytical System of Chiplets for AI Devices at the Edge) is a design space exploration framework for heterogeneous chiplet-based System-in-Package architectures for edge ML.

CASCADE overview: a chiplet menu feeds design space exploration that produces a System-in-Package recipe card

DreamRAM

DreamRAM is a configurable modeling and design space exploration tool for custom 3D die-stacked DRAM architectures based on HBM.

DreamRAM framework: input parameters, modeling flow from floorplanning to energy, and output metrics

McPAT CPU Models

We have developed McPAT power models for a recent high-performance multicore CPU.

Bar chart of core area for McPAT model revisions compared with the measured POWER7 subset

NVMExplorer

NVMExplorer is a cross-stack design space exploration framework for evaluating and comparing on-chip memory solutions including emerging, embedded non-volatile memories.

NVMExplorer

ML Systems & Recommendation

DeepRecSys

DeepRecSys provides an end-to-end infrastructure to study and optimize at-scale neural recommendation inference.

General architecture of neural personalized recommendation models with dense and sparse features

RecPipe

RecPipe provides an end-to-end system to study and jointly optimize recommendation models and hardware for at-scale inference.

End-to-end RecPipe system diagram for multi-stage recommendation across application, algorithm and hardware

Reliability & Resilience

Ares

Ares is a framework for quantifying the resilience of deep neural networks.

Ares fault-injection framework applied to an example DNN, with weight, activation and state fault models

GoldenEye

GoldenEye is a functional simulator with fault injection capabilities for common and emerging numerical formats, implemented for the PyTorch deep learning framework.

Number formats explored with GoldenEye: FP32, fixed point, INT8, block floating point and AdaptivFloat

Benchmarks & Workloads

BayesSuite

BayesSuite is a collection of Bayesian inference workloads written in Stan framework.

BayesSuite logo

Fathom

Fathom is a collection of workloads for benchmarking modern machine learning techniques.

Fathom logo: Reference Workloads for Modern Deep Learning

MachSuite

MachSuite is a benchmark suite for high-level synthesis and accelerator-centric architectures.

Instruction mix (compute, memory, branch) across the MachSuite benchmarks

Workload Characterization & Analysis

CHAMPVis

CHAMPVis, Comparative Hierarchical Analysis of Microarchitectural Performance Visualization.

CHAMPVis views comparing microarchitectural performance across applications using the TopDown hierarchy

LLVM-Tracer

LLVM-Tracer is an LLVM instrumentation pass to print out a dynamic LLVM IR execution trace, including dynamic values and memory addresses.

A C loop and the dynamic LLVM IR trace generated from it

WIICA

WIICA is the Workload ISA Independent Characterization for Applications tool.

Instruction breakdown for x86, LLVM IR and simplified IR across SPEC benchmarks

Compilers & CPU Simulation

ILDJIT

ILDJIT is our compilation framework using a high-level intermediate representation.

HELIX compilation infrastructure built on ILDJIT: static compiler, generated threads and a multicore target

XIOSim

XIOSim is a multicore x86 performance simulator that models in-order and out-of-order cores with RingCache, developed as part of the HELIX automatic parallelization project.

XIOSim execution model: a workload under the Pin VM drives simulated cores, uncore and a power model