The workshop
All projects (8)
01C++ header-only library for automatic tensor kernel optimization for CPU inference. Uses expression templates and metaprogramming for zero-overhead abstraction, automatic kernel fusion, and compile-time optimization. Supports AVX2/AVX-512/NEON.
02AI-driven system that uses LLMs to iteratively optimize C++ kernel performance. Orchestrates compilation, correctness validation with GoogleTest, and performance measurement with Google Benchmark in an automated loop.
03GPU-accelerated sparse finite volume solver for hyperbolic PDEs on complex 2D geometries. Uses a novel Compressed Sparse Row of intervals data structure. Built on Kokkos for portable CPU/GPU performance.
04Contributor to a C++ library for adaptive mesh refinement (AMR) using interval-based set algebra. Supports patch-based, cell-based, and multiresolution methods from a single data structure.
C++AMRHPCNumerical Methods
05Python bindings for the Samurai AMR library using pybind11. Standalone package with Meson build system, MPI support, and conda integration for rapid prototyping and visualization.
06Runs Google Benchmark binaries with statistical meta-repetitions. Automatically re-runs only unstable benchmarks until confidence intervals converge. Pip-installable.
PythonBenchmarkingStatistics
07Visual tool for detecting performance regressions from Google Benchmark traces. Provides severity-based coloring, per-kernel aggregation, and CI mode with configurable thresholds.
08Optimized pow() implementations in C++ achieving 6.6x speedup over std::pow for integer exponents using hierarchical exponentiation. Accuracy within 1-2 ULP.
C++PerformanceNumerical Computing