3 comments

  • vorticalbox 2 hours ago
    This reminds me of https://dnhkng.github.io/posts/rys/

    David looks into the LLM finds the thinking layers and cut duplicates then and put them back to back.

    This increases the LLM scores with basically no over head.

    Very interesting read.

    • renticulous 8 minutes ago
      Jeff Dean says models hallucinate because their training data is "squishy."

      But what's in the context window is sharp, the exact text or video frame right in front of them.

      The goal is to bring more of the world into that context.

      Compression gives it intuition. Context gives it precision.

      Imagine if we could extract the model's reasoning core and plug it anywhere we want.

  • l4tq3 1 hour ago
    [dead]
  • 34ylsh 1 hour ago
    [flagged]