back
user profile

convexstrictly

1,147karma·112submissions·March 26, 2023
recent activity (112 total)
comment
Pricing input: $8/1M tokens output: $24/1M tokens https://docs.mistral.ai/platform/pricing/ …
2y ago·view thread
comment
Jeremy makes compelling arguments. Here are some more mundane corollaries: It is only a matter of time before you and your company are affected by the pending regulations. In the future, almost all …
2y ago·view thread
comment
Some information here. https://www.ntia.gov/federal-register-notice/2024/dual-use-f... …
2y ago·view thread
comment
The comments will inform the drafting of regulations on open weight models under the Biden executive order on AI using his powers under the Defense Production Act. Fact Sheet: https://www…
2y ago·view thread
comment
"By enabling the use of a single high-precision base model accompanied by multiple 1-bit deltas, BitDelta dramatically reduces GPU memory requirements by more than 10x, which can also be translat…
2y ago·view thread
comment
https://github.com/FasterDecoding/BitDelta
2y ago·view thread
comment
Twitter summary: https://twitter.com/ssgrn/status/1738256456250470853 Github: https://github.com/KaiNylund/lm-weights-encode-time …
2y ago·view thread
comment
Everything you say makes sense. Training is definitely more compute intensive than inference. Training is both memory throughput and compute constrained. Much research in speeding up training goes i…
2y ago·view thread
comment
Could you explain the blockers to getting back-propagation working well on your chips?
2y ago·view thread
comment
Research suggesting that much of the power of the transformer architecture comes from associative recall over long sequences that does not require scaling model dimensions. They design state space mo…
2y ago·view thread
comment
"... we find that a duo of a 1.3B generation model and a 1.3B verifier model can achieve 81.5% accuracy, outperforming existing models that are orders of magnitude larger."
2y ago·view thread
comment
The paper claims it builds upon the concepts in HashGraph, an efficient CUDA hashtable implementation. HashGraph (2019) https://arxiv.org/abs/1907.02900 Anyone know what the most…
2y ago·view thread
comment
Github repo https://github.com/harp-lab/gdlog
2y ago·view thread
comment
Twitter thread with video introduction https://twitter.com/1x_tech/status/1730610445541638378 …
2y ago·view thread
comment
Emily Chang from Bloomberg reports: Satya Nadella was "blindsided" and is furious. Details in paywalled article: https://www.bloomberg.com/news/articles/2023-11-18&…
2y ago·view thread
comment
If you were logged into Bing, those prompts may be in your history. They can be viewed using the Edge browser. A few weeks ago, I had spotty service with Bing Chat where it would keep resetting the c…
2y ago·view thread
comment
Github Copilot Chat is in beta and is marketed as a separate product. I suspect it is using a later generation (and better) underlying model. I used it many months ago, and I agree it is much better …
2y ago·view thread
comment
Yes, I have been using it the last few days. I haven't noticed the failure case of Bing Chat not answering the question at all. Are you sure you had the modes I recommended turned on? Try rep…
2y ago·view thread
comment
After the recent updates to the UI, the old 50 messages every 3 hours warning went away. Though I haven't really tried pushing it past that limit to tell if that restriction is still there.
2y ago·view thread
comment
Bing Chat (Choose "Creative Mode" or "Use GPT-4" depending on UI) is GPT-4 with retrieval augmented generation using the Bing Search engine. It seems to be tuned a little differen…
2y ago·view thread
comment
"Greg Brockman works 60 to 100 hours per week, and spends around 80% of the time coding. Former colleagues have described him as the hardest-working person at OpenAI." https://tim…
2y ago·view thread
comment
"I am currently on leave from MIT and spending it at OpenAI." https://madry.mit.edu/
2y ago·view thread
comment
Jakub Pachocki and Szymon Sidor have worked on mu-parametrization/tensor programs and Dota 2. https://www.semanticscholar.org/author/J.-Pachocki/2713380?s... As @eachro…
2y ago·view thread