Which one is more important: more parameters or more computation? (2021)

(parl.ai)

30 points | by jxmorris12 1 day ago

3 comments

vorticalbox 2 hours ago
This reminds me of https://dnhkng.github.io/posts/rys/
David looks into the LLM finds the thinking layers and cut duplicates then and put them back to back.
This increases the LLM scores with basically no over head.
Very interesting read.
[-]
- renticulous 8 minutes ago
  Jeff Dean says models hallucinate because their training data is "squishy."
  But what's in the context window is sharp, the exact text or video frame right in front of them.
  The goal is to bring more of the world into that context.
  Compression gives it intuition. Context gives it precision.
  Imagine if we could extract the model's reasoning core and plug it anywhere we want.
l4tq3 1 hour ago
[dead]
34ylsh 1 hour ago
[flagged]