I feel like Debug should be lazily emitted altogether when it’s first used - just a special marker that’s never expanded since 99% of the Debug implementations aren’t used and having the rest marked #[cold] as inline is obviously wrong. Of course implementing it in practice sounds exceptionally difficult.
That being said, even the justifying performance improvement PR was itself a mix of improvements and regressions
Lazily emitted from what? You would need information about the type at runtime to derive the implementation, and normally that information is not available at runtime. The debug implementation itself is probably within spitting distance of any other representation that would be sufficient to lazily emit the debug implementation.
You can do it at compile time. Or more likely, link time. If the implementation is ever actually used, it is kept in the binary. Otherwise it is used. The compiler can do the same, keeping all derives as just markers until it finds a place that actually uses it then firing off a background worker to compile the derive impl.
I’m fairly convinced that Debug should never be inlined. Display probably neither, the fmt machinery is heavy enough that not inlining is probably not a bottleneck even in serialization-heavy workloads. I’ve had to #[inline(never)] some of my own Debug/Display impls, shrinking the binary by tens of kilobytes (out of a few hundred, so relatively a significant reduction).
For what it's worth, according to the PR that added the annotation [0] doing so generally resulted in decreases in compile times and binary sizes on benchmarks. Furthermore, an additional experiment that avoided emitting the inline attribute on structs with >5 fields resulted in benchmark regressions compared to always emitting the attribute [1]. I'd guess this is one of those things which may help in aggregate but hurts for specific cases.
That being said, one of the Rust devs indicated in the corresponding lobste.rs discussion [2] that they're open to revisiting/rebalancing things if they get enough bug reports indicating something is up, so it might not hurt to tag onto the bug report the author will (hopefully) eventually submit.
A reasonable conjecture was raised on lobsters that this is because `#[inline]` makes actual codegen (LLVM IR and down from MIR) lazy, and most `Debug` impls are never used.
How much code is never used and compilation could be skipped entirely? Maybe applying a reachability pass to skip compiling unused code would be helpful.
A cross-crate dead code analysis would mean that compilation of a crate now depends on information about its dependents. This would break reuse of compiled crates and cause recompiles when the analysis changes.
Something does seem a little off about this, though. Ideally for this `Debug` case there would be an annotation that says "compile this lazily, don't inline". Maybe there doesn't even need to be a new annotation, just `#[inline] #[cold]`. Which looks pretty weird, but might work already.
I wonder if they ever considered a table-driven approach for `#[derive(Debug)]`, as `facet` [1] does. That would have been my first instinct for something this formulaic where binary size and compilation time matter more than execution speed. But my impression is facet hasn't quite realized its promise on those fronts, so maybe the table-driven approach in std was similarly tried and rejected.
Or just try to avoid all of these optimization guesses by using Profile-Guided Optimization (PGO), that inserts/deletes all inlines based on actual application runtime profile.
The codebase in question (uv) uses PGO already. I suspect there isn’t a general way to guarantee that PGO ensures that only the “right” things get inlined.
There's a joke that the LLVM heuristic for whether to inline a function is "return true;" LLVM tends to inline aggressively.
You can control this behavior with opt level "s" or "z" or "#[inline(never)]", but be aware that too little inlining can have large negative performance impacts.
It's hard to get inlining exactly right without profile guided optimization.
That being said, even the justifying performance improvement PR was itself a mix of improvements and regressions
That being said, one of the Rust devs indicated in the corresponding lobste.rs discussion [2] that they're open to revisiting/rebalancing things if they get enough bug reports indicating something is up, so it might not hurt to tag onto the bug report the author will (hopefully) eventually submit.
[0]: https://github.com/rust-lang/rust/pull/117727
[1]: https://github.com/rust-lang/rust/pull/118031
[2]: https://lobste.rs/s/dldhpw/rust_s_derive_often_implies_inlin...
Something does seem a little off about this, though. Ideally for this `Debug` case there would be an annotation that says "compile this lazily, don't inline". Maybe there doesn't even need to be a new annotation, just `#[inline] #[cold]`. Which looks pretty weird, but might work already.
[1] https://crates.io/crates/facet
You can control this behavior with opt level "s" or "z" or "#[inline(never)]", but be aware that too little inlining can have large negative performance impacts.
It's hard to get inlining exactly right without profile guided optimization.