hn
.
reader
top
new
show
ask
jobs
back
MA
user profile
matt_d
21,167
karma
·
3,275
submissions
·
April 21, 2014
recent activity
(3,275 total)
2 pts
What "Memory Compiler" Actually Means: From Bitcells to GDS Tiling
(thecloudlet.github.io)
2mo ago
·
discuss
6 pts
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-Offs, and Perf
(arxiv.org)
2mo ago
·
discuss
2 pts
MIT EECS/CSAIL Agentic Coding in Practice Seminar Series
(people.csail.mit.edu)
2mo ago
·
discuss
1 pts
The Load-Balance Problem Behind Hybrid Parallelism
(hecate0821.github.io)
2mo ago
·
discuss
2 pts
Delayed Tensor Parallelism for Faster Transformer Inference
(blog.kog.ai)
2mo ago
·
discuss
2 pts
Continuous Diffusion Models Can Obey Formal Syntax
(arxiv.org)
2mo ago
·
discuss
2 pts
HartBreaker: Deterministic Fuzzing of Multi-Hart RISC-V CPUs
(comsec.ethz.ch)
2mo ago
·
discuss
10 pts
Tuning LLVM's SLP Vectorizer Cost Model
(blog.kaving.me)
2mo ago
·
discuss
2 pts
A Friendly Tour of Substructural, Uniqueness, Ownership, Capabilities and more!
(federicobruzzone.github.io)
2mo ago
·
discuss
1 pts
FlashLib: Bringing Flash Magic to Classical Machine Learning Operators
(flashml-org.github.io)
2mo ago
·
discuss
1 pts
FML-Bench: A Controlled Study of AI Research Agent Strategies
(arxiv.org)
2mo ago
·
discuss
2 pts
Finding deadlocks in CuTe kernels with SPIN
(metaworld.me)
2mo ago
·
discuss
2 pts
A Case for Tracing Based DSL Kernel Languages
(metaworld.me)
2mo ago
·
discuss
4 pts
You don't need all the LLM benchmarks
(alex.smola.org)
2mo ago
·
discuss
1 pts
Elusive order of async GPU kernels: scheduling, abstractions, DSL implications
(ianbarber.blog)
2mo ago
·
discuss
1 pts
MileStone: A Multi-Objective Compiler Phase Ordering Framework
(arxiv.org)
2mo ago
·
discuss
4 pts
SSV: Sparse Speculative Verification for Efficient LLM Inference
(arxiv.org)
2mo ago
·
discuss
2 pts
Characterizing Real-World Bugs in Tile Programs for Automated Bug Detection
(arxiv.org)
2mo ago
·
discuss
3 pts
Characterization of machine learning compilers for LLM inference on NVIDIA GPUs
(link.springer.com)
2mo ago
·
discuss
2 pts
Chip design from the bottom up – Reiner Pope [video]
(youtube.com)
2mo ago
·
discuss
2 pts
LT2: Linear-Time Looped Transformers
(charlesdddd.github.io)
2mo ago
·
discuss
6 pts
Event Tensor: A Unified Abstraction for Compiling Dynamic Megakernel
(arxiv.org)
2mo ago
·
discuss
1 pts
PopPy: Opportunistically Exploiting Parallelism in Python Compound AI Apps
(arxiv.org)
2mo ago
·
discuss
105 pts
CODA: Rewriting Transformer Blocks as GEMM-Epilogue Programs
(arxiv.org)
2mo ago
·
12 comments
6 pts
[RFC] Open Access to Standards Documents – LLVM Project
(discourse.llvm.org)
2mo ago
·
discuss
6 pts
Curly braces: An evolution of UNIX and C
(thalia.dev)
2mo ago
·
2 comments
2 pts
NanoTag: Systems Support for Efficient Byte-Granular Overflow Detection on Arm
(github.com)
2mo ago
·
discuss
2 pts
InferenceBench: A Benchmark for Open-Ended Inference Optimization by AI Agents
(inferencebench.ai)
2mo ago
·
discuss
← prev
1
…
8
9
10
11
12
…
110
next →