back
user profile
convexstrictly
1,147karma·112submissions·March 26, 2023
recent activity (112 total)
comment
"We start from the surprising finding that certain attention heads have a clear activation distribution difference for true and false statements. Probing at these points yields upwards of 83% acc…
comment
"Orca surpasses ... Vicuna-13B by more than 100% in complex zero-shot reasoning benchmarks like Big-Bench Hard (BBH) and 42% on AGIEval. ... reaches parity with ChatGPT on the BBH benchmark and s…
comment
"We present SpQR, which allows lossless LLM inference at 4.75 bits with a 15% speedup. You can run a 33B LLM on a single 24GB GPU fully lossless. SpQR works by isolating sensitive weights with hi…
comment
The Brainformer building block is designed using neural architecture search. "Brainformer consistently outperforms the state-of-the-art dense and sparse Transformers, in terms of both quality and…
comment
Relevant codebase is https://github.com/getcursor/human-eval
comment
They claim average 1.45x speedup and maximum 1.7x speedup over GPTQ. High-level idea: About 1% of the weights contribute greatly to quantization error. So skip the quantization of these weights.
comment
For fast inference, the HuggingFace cofounder, Thom Wolf recommends their text-generation-inference library https://github.com/huggingface/text-generation-inference https:/…
comment
Agree. The Stable Diffusion Open RAIL M license: "You agree not to use the Model or Derivatives of the Model ... To defame, disparage or otherwise harass others" Does "disparage" …
comment
Research in the LLM space is moving so fast, I doubt any architecture will "stick".
comment
Do you consider the Unity model "sleazy"? But I agree that it was poor form to call it open-source at the time. Sadly, it seems to be standard practice in the LLM space to release models wi…
comment
Yann LeCun: "No. But it's not because we don't want to. It's because of complicated legal issues." https://twitter.com/ylecun/status/1651782621540524…
comment
https://huggingface.co/tiiuae/falcon-40b
comment
Closed beta runs for a week. 4 bit floats. Quantization on top of quantization. Finetune Alpaca 7B in ~3 hours on an A40 GPU/RTX3090. Paging of optimizer states. Uses LoRA.
Integrated with H…
comment
Video: https://www.youtube.com/watch?v=iI8e3kU11eU