hn
.
reader
top
new
show
ask
jobs
back
MA
user profile
matt_d
21,168
karma
·
3,275
submissions
·
April 21, 2014
recent activity
(3,275 total)
1 pts
Tessera: Unlocking Heterogeneous GPUs Through Kernel-Granularity Disaggregation
(arxiv.org)
4mo ago
·
discuss
2 pts
From SIMT to Systolic: A Foundation for GPU and TPU Architecture
(twitter.com)
4mo ago
·
discuss
2 pts
Packrat Parsing at the Speed of Wasm [video]
(youtube.com)
4mo ago
·
discuss
3 pts
Sparser, Faster, Lighter Transformer Language Models
(arxiv.org)
4mo ago
·
discuss
1 pts
When GPUs Fail Quietly: Observability-Aware Early Warning Beyond Telemetry
(arxiv.org)
4mo ago
·
discuss
3 pts
Stupid RCU Tricks: Corner-Case RCU Implementations
(people.kernel.org)
4mo ago
·
discuss
1 pts
How Many Compilers Is Too Many? V8's History, Tradeoffs, and Architecture [video]
(youtube.com)
4mo ago
·
discuss
1 pts
Fully-Automatic Type Inference for Borrows with Lifetimes
(al.radbox.org)
4mo ago
·
discuss
119 pts
The GNU libc atanh is correctly rounded
(inria.hal.science)
4mo ago
·
29 comments
1 pts
MMU Handbook: Memory Management Units and TLBs
(kalairajah-personal.github.io)
4mo ago
·
discuss
2 pts
PEP 831 – Frame Pointers Everywhere: Enabling System-Level Observability
(peps.python.org)
4mo ago
·
discuss
13 pts
UpDown: Efficient Manycore based on Many Threading & Scalable Memory Parallelism
(people.cs.uchicago.edu)
4mo ago
·
1 comments
146 pts
Reflections on 30 years of HPC programming
(chapel-lang.org)
4mo ago
·
123 comments
1 pts
Recent lld/ELF performance improvements
(maskray.me)
4mo ago
·
discuss
45 pts
Circuit Transformations, Loop Fusion, and Inductive Proof
(natetyoung.github.io)
4mo ago
·
3 comments
2 pts
AI for Systems: Using LLMs to Optimize Database Query Execution
(together.ai)
4mo ago
·
discuss
3 pts
Optimization of 32-bit Unsigned Division by Constants on 64-bit Targets
(arxiv.org)
4mo ago
·
discuss
4 pts
GCC Translation Validation Part 6: Uninitialized Memory
(kristerw.github.io)
4mo ago
·
discuss
2 pts
Agentic Code Optimization via Compiler-LLM Cooperation
(arxiv.org)
4mo ago
·
discuss
1 pts
Bespoke OLAP: Using AI to Synthesize Workload-Specific DBMS Engines from Scratch
(ucbskyadrs.github.io)
4mo ago
·
discuss
2 pts
Understanding Agents: Code Coverage for Coding Agents
(blog.asymmetric.re)
4mo ago
·
discuss
1 pts
Cyclotron: The Streaming Multiprocessor Abstraction Is Broken [pdf]
(capra.cs.cornell.edu)
4mo ago
·
discuss
2 pts
UCCL-EP: Portable Expert-Parallel Communication – Full Results
(uccl-project.github.io)
4mo ago
·
discuss
4 pts
vLLM IR: A Functional Intermediate Representation for vLLM
(github.com)
4mo ago
·
discuss
1 pts
Test-Time Scaling Makes Overtraining Compute-Optimal
(arxiv.org)
4mo ago
·
discuss
1 pts
Analyzing Reverse Address Translation Overheads in Multi-GPU Scale-Up Pods
(arxiv.org)
4mo ago
·
discuss
3 pts
Tracing a Full MoE Training Step Through the XLA Compiler
(patricktoulme.substack.com)
4mo ago
·
discuss
2 pts
Breaking Down the Cerebras Wafer Scale Engine
(wafer.substack.com)
4mo ago
·
discuss
2 pts
The need for better compiler frontend benchmarks: Carbon's benchmarking approach
(discourse.llvm.org)
4mo ago
·
discuss
3 pts
DAXFS: A Lock-Free Shared Filesystem for CXL Disaggregated Memory
(arxiv.org)
4mo ago
·
discuss
← prev
1
…
13
14
15
16
17
…
110
next →