back
user profile

convexstrictly

1,147karma·112submissions·March 26, 2023
recent activity (112 total)
comment
"We start from the surprising finding that certain attention heads have a clear activation distribution difference for true and false statements. Probing at these points yields upwards of 83% acc…
3y ago·view thread
comment
"Orca surpasses ... Vicuna-13B by more than 100% in complex zero-shot reasoning benchmarks like Big-Bench Hard (BBH) and 42% on AGIEval. ... reaches parity with ChatGPT on the BBH benchmark and s…
3y ago·view thread
comment
"We present SpQR, which allows lossless LLM inference at 4.75 bits with a 15% speedup. You can run a 33B LLM on a single 24GB GPU fully lossless. SpQR works by isolating sensitive weights with hi…
3y ago·view thread
comment
The Brainformer building block is designed using neural architecture search. "Brainformer consistently outperforms the state-of-the-art dense and sparse Transformers, in terms of both quality and…
3y ago·view thread
comment
Relevant codebase is https://github.com/getcursor/human-eval
3y ago·view thread
comment
They claim average 1.45x speedup and maximum 1.7x speedup over GPTQ. High-level idea: About 1% of the weights contribute greatly to quantization error. So skip the quantization of these weights.
3y ago·view thread
comment
For fast inference, the HuggingFace cofounder, Thom Wolf recommends their text-generation-inference library https://github.com/huggingface/text-generation-inference https:/…
3y ago·view thread
comment
Agree. The Stable Diffusion Open RAIL M license: "You agree not to use the Model or Derivatives of the Model ... To defame, disparage or otherwise harass others" Does "disparage" …
3y ago·view thread
comment
Research in the LLM space is moving so fast, I doubt any architecture will "stick".
3y ago·view thread
comment
Do you consider the Unity model "sleazy"? But I agree that it was poor form to call it open-source at the time. Sadly, it seems to be standard practice in the LLM space to release models wi…
3y ago·view thread
comment
Yann LeCun: "No. But it's not because we don't want to. It's because of complicated legal issues." https://twitter.com/ylecun/status/1651782621540524…
3y ago·view thread
comment
https://huggingface.co/tiiuae/falcon-40b
3y ago·view thread
comment
Closed beta runs for a week. 4 bit floats. Quantization on top of quantization. Finetune Alpaca 7B in ~3 hours on an A40 GPU/RTX3090. Paging of optimizer states. Uses LoRA. Integrated with H…
3y ago·view thread
comment
Video: https://www.youtube.com/watch?v=iI8e3kU11eU
3y ago·view thread