Fwiw on the compiler front, I used to work on the msvc tool chain. It has been doing lto (aka ltcg) for a very long time at reasonably good throughput. That tool chain also has done incremental compilation on a per function basis for about 7 years. Sharing header parsing across translation units has been possible through PCH for decades. Even without PCH, inline header functions only get codegened once in ltcg mode (not counting inline expansion).
In a large multi-dll build like Windows or Office, the import/export information across modules is computed without doing codegen, so ltcg codegen can be done in parallel regardless of the module dependency graph.
I never worked deeply in the build systems for Linux based OSes or other unixes, but I gather that some symbol visibility choices make the build situation worse in Unix systems than Windows.
In a large multi-dll build like Windows or Office, the import/export information across modules is computed without doing codegen, so ltcg codegen can be done in parallel regardless of the module dependency graph.
I never worked deeply in the build systems for Linux based OSes or other unixes, but I gather that some symbol visibility choices make the build situation worse in Unix systems than Windows.