22 comments

  • nneonneo 14 hours ago
    At this point, in the field of AI research:

    - papers are written by AI (as pointed out in this article, and as obvious to anyone who spends a while actually reading recent AI research)

    - papers are reviewed by AI (NeurIPS is doing an AI assisted review experiment - https://neurips.cc/Conferences/2026/ai-reviewing-experiment - and I feel the trend is moving towards AI reviewers whether we like it or not)

    - papers are read, summarized and digested by AI, because there are just so many papers at leading AI conferences that nobody has time to eyeball them all

    We are very rapidly automating humans out of the academic publication loop here.

    • jsrozner 13 hours ago
      Idk if we're automating humans out of the publishing loop as much as rapidly automating the production of crap. I had a very similar experience reviewing for EMNLP recently.

      We are nowhere near AI being able to judge the quality of research (in fact, one might reasonably state that even most humans can't really judge the quality of research). Most things in society are not like math: we can't automate (via verification) our way out of noise overwhelming the signal.

      Folks are willing to entirely abuse the public resource that is faithful, honest reviewing. (This is unsurprising; the abuse of the commons / public resources has been rising for a long time). There isn't a good solution other than something akin to draconian social scoring to limit access to the reviewing system.

      • boplicity 13 hours ago
        > draconian social scoring

        This is a legitimate question: should people have reputations? Should their behavior be made more visible publicly, both good and bad? How?

        In a small contained society, where consequences are more directly affecting individuals immedately, these questions don't need to be asked, because they're inherantly answered. We now are a society with billions of people, and dire consequences sometimes deferred for a generation, or more. Part of our general failing is the lack of good answers to the above questions. For many people, there are rarely negative consequences for causing harm to others, and the rewards can be very great indeed.

        • ryandrake 12 hours ago
          "Consequences and reputations stemming from one's actions" isn't necessarily draconian social scoring. Without a structure for imposing consequences on wrongdoers, we're not a society, we're just monkeys flinging poo at each other.
          • Eddy_Viscosity2 55 minutes ago
            Even monkeys have society and they tend to only throw poo at those who have done wrong by that society. Monkeys have and try to maintain reputations.
          • blackqueeriroh 11 hours ago
            The problem is that everyone imagines some sort of just and fair arbiter of these things, when the reality is all of the social scoring and consequences and reputations rarely actually stem from one’s actions, and far more often stem from how much money someone has put in someone’s pocket, who someone knows, the color of someone’s skin, or what’s between someone’s legs.

            Until we’re actually serious about treating people equitably, (not equally, as that would simply leave the lopsided power structure we have in place) we aren’t getting out of this.

            • somenameforme 9 hours ago
              I don't understand the desire to argue against this in identity politics terms, which I assume you realize is going to be extremely divisive. There's a far more simple argument that people, independent of ideological framing, would tend to agree upon. And that's what is valuable to one person isn't valuable, to the same degree, to another. For things that are completely illegal, like murder, such a system would be pretty useless because we already have systems in place to punish that sort of behavior and defacto social scoring on it as well.

              So this is going to come down to things that aren't usually illegal or 'that illegal', but otherwise affect society. So for instance one group might want to reasonably punish the directors of a company that causes emissions. Another group might want to reasonable reward the directors of a company that creates a large number of desirable jobs. And even in this one example you immediately end up with a weird scenario. Is it okay to pollute as long as you make enough jobs? Who gets to decide that?

              There's no need for cartoon villainy for this to be a bad idea. Even completely well intentioned, it just doesn't work out well. And the more diverse a society, the worse it's going to work.

              • knollimar 2 hours ago
                How do you combat review bombing or someone just being bigoted?
            • bell-cot 4 hours ago
              > The problem is that everyone imagines ...

              I'd phrase it "most people want to imagine". Or will claim they want to - since not believing sounds depressing, or suggests that the person is an evildoer hoping to escape justice. (But unfortunately, people usually disagree about exactly what would be "just" or "fair". While wanting to imagine that they don't. Yes, the problem just went meta.)

              > when the reality is ...

              Try asking some really old folks about how much drama, inequity, and nastiness there often was in social settings where everyone was the same color and gender, nobody was notably wealthy, and nobody had any great connections. Or talk to an experienced junior high teacher. Or read some history. Humans are quite capable of dividing themselves into camps over any "differences" that they're able to perceive. Or invent, since dividing themselves into camps is often the unspoken objective.

          • dwattttt 11 hours ago
            For a more mathematically rigorous treatment, negative feedback is fundamental to Control Theory; it allows us to stabilise processes that would otherwise go out of control.
          • csallen 11 hours ago
            People react negatively to this, because they fear a dystopian society of the sort we've seen in plenty of movies, and rightfully so.

            But it's also worth pointing out that "consequences and reputations stemming from one's actions" is already the world we live in and always have lived in. Hell, even Hacker News has karma points, downvoting, shadow banning, and the like. There's no such thing as a society with zero consequences and zero reputations. The only real question is a matter of degree, structure, severity, reach, and various idiosyncrasies that differ across cultures.

            So it would be nice to have a more nuanced discussion about this instead of treating it like a 0 or 1 decision.

          • globalnode 4 hours ago
            and then you have to ask who defines good and bad behaviour and who determines consequences when power is lop sided.
          • cindyllm 12 hours ago
            [dead]
        • BobbyTables2 11 hours ago
          Isn’t China doing something like this ?

          Wonder if it could degrade to something worse than the status quo where reputations are falsely tarnished by competitors purely as a weapon to get ahead…

          • nneonneo 10 hours ago
            The "Chinese social credit" system is vastly overblown.

            I'm living here (China, Beijing and Jiangsu) now, and there are exactly two cases in the past eight months where the "social credit" system has had any impact at all:

            - To open "take first pay after" vending machines. These are vending machines which are basically big locked fridges; if your credit score is high enough, you can open the machines, take whatever drinks/snacks you want, and they'll use (I presume) computer vision to charge you afterwards. If you don't have a high enough credit score, tough, you can't use them. (But there's almost always normal pay-first vending machines nearby).

            - To borrow mobile charging packs (powerbanks). Some operators will let you "swipe your credit" (check your credit score) to take one without paying a deposit. If your credit score isn't high enough, you first pay e.g. a ¥99 deposit (~$15USD) which gets returned when you return the powerbank, not a big deal.

            That's...it. My credit score is high enough on only one platform (Alipay), so I get to try what happens with both "high" and "low" credit, and I can confidently say these are the *only* two cases where I have even been asked to show the credit score. I have taken multiple train trips, bought lots of stuff in stores and restaurants, etc. without ever touching the "social credit score".

            P.S. I literally don't even know how to raise my credit score, and neither do many of the folks who live here - again, not that it matters, because the score is simply not that important.

            • inigyou 4 hours ago
              But the USA has a social credit score that affects a lot more about your life, doesn't it? Like whether you're allowed to have a house or a car.
              • ben_w 4 hours ago
                Social credit? Or money credit?

                (Genuine question, I'm not American and don't have any desire to move to the US).

                • xoa 2 hours ago
                  >Social credit? Or money credit?

                  Neither. Above poster is wrong. You don't require anything to own a house or car in the US except your own cash/equiv. If you wish to borrow other people's money or property (in some cases at all when counterparty risk is high enough or buyers competitive enough, in most cases it's more a matter of interest rate), a sufficient credit score is the most standard and easy way to "make the case" that you will repay. But that's truly not at all required. I bought my current truck for $8100 cash on the table, I just showed up at the other guy's place and inspected it and chatted, looked good enough, paid him and we did the title change document, I screwed on a new plate and off I drove with it. Effectively zero dealers will not accept cash upfront. You can buy land and just build your own place.

                  Obviously, a lot of us find having credit very useful. And it's also both convenient and good for fairness/economic velocity/efficiency not to generally have to explicitly put up any extra collateral to get it or go through some complex old fashioned social dance. Hence "credit scores" that systemize a lot. But claiming they decide "whether you're allowed to have a house or a car" warps things badly.

                • inigyou 2 hours ago
                  They are the same thing. The Chinese social credit is mostly about money. The USA money credit is affected by social factors. They are the same.
            • rogerrogerr 7 hours ago
              I remember in the west a few years ago we were seeing purported pictures of electronic billboards in China shaming low-social-credit-havers. Was that all a weird psyop?
              • scns 5 hours ago
                > electronic billboards in China shaming low-social-credit-havers

                It is people who did not pay their debts. Dr Jonathan Tam has a video on YT explaining it called:

                Why the World Fell for China's Fake Dystopia

                https://www.youtube.com/watch?v=Ecx45eOuc0E

        • gspr 8 hours ago
          There's a huge difference between a completely decentralized reputation system like the interpersonal one that has existed since humans first became sentient, and an artificially built top-down reputation system run by a powerful entity (be it the state like in China or megacorps like in the West).
        • attila-lendvai 12 hours ago
          in short: our social sructure has built out so far beyond the Dunbar's number that it has become a haven for psychopaths.
        • onetokeoverthe 12 hours ago
          [dead]
      • ludicrousdispla 6 hours ago
        The 'named' reviewers typically pass on the review responsibilty to grad students or less-senior colleagues. This is currently a generally accepted and sanctioned/encouraged practice. A social scoring system would need to eliminate that for true accountability.
        • gradstudent 4 hours ago
          Reviews are typically in conjunction with junior colleagues and the senior person is involved in the discussion and held responsible for the outcome. The junior folks then end up in the proceedings as external reviewers. In my experience anyway. I don't see an issue here. How else do you train junior researchers to review? Its how I learned and all those around me as well.
        • weiliddat 6 hours ago
          Yeah it's a well-known secret, and I think any sort of scoring system misses the point. The scentific method (and system, incl. peer reviews) is a public good (like good governance) that requires moral acknowledgement, the necessary social/cultural/political/academic pressures, and discipline of taking care of it by participating parties.

          If you add a score/metric this incentivizes the wrong thing, like how money, funding, and career already incentivizes all the wrong non-scientific behaviors (like faking data, p-hacking, etc.)

          • StableAlkyne 19 minutes ago
            > If you add a score/metric this incentivizes the wrong thing,

            Don't we already have all of these things because academia decided to use H-index as a metric for career impact?

            • weiliddat 1 minute ago
              Yeah exactly. One more metric isn't going to solve it.
      • azan_ 12 hours ago
        People should get paid for reviewing. That’s the solution. Publishing companies rack in billions of dollars in pure profit exploring free labor. Once people actually get paid for reviewing, it becomes much faster and higher quality and you won’t need AI triage.
        • Sharlin 10 hours ago
          Doesn’t work anymore because people would just have ChatGPT generate their reviews and get paid for doing nothing.
          • zx8080 10 hours ago
            So in sum:

            - it did not work before because greed

            - will not work anymore because AI

            Is human science dead?

            • StableAlkyne 9 minutes ago
              Nah, people will just come up with different (probably broader) heuristics when reviewing.

              We have a past example of this in the US, too. Certain countries have historically had a huge problem with paper mills. Because of this, most people in the US/EU do have a negative bias once they see the author affiliations of this countries. Yes, it isn't fair to researchers who aren't pulling some shenaniganry. But it is a common shortcut; the sixth or seventh time you've spent a few hours reviewing a paper filled with bullshit, you probably would develop it too.

              I imagine we will see similar things come up: the most obvious I can think of is, if it's not a well-known institution then it might be bullshit.

              Hell, we are already seeing something similar in FOSS, where many high profile projects have completely banned AI contributions due to all the low effort slop PRs.

            • inigyou 1 hour ago
              Science as an activity will never die, but science as an institution is in big trouble.
          • afdbcreid 10 hours ago
            Then perhaps people need to pay to get their paper reviewed?
            • azan_ 7 hours ago
              People already do.
        • tollgategit 6 hours ago
          Alternatively, put the pressure on the incoming requests for review.

          Pay a deposit to have your work reviewed, if it's accepted as a good faith submission, the deposit is returned minus a minimal, irrelevant fee, if it is deemed in bad faith, take all of the deposit as a time-waste tax. With enough cases going wrong, the flood of slop slows down and maintainers have to pull the trigger as often.

          This was the solution several people had proposed for Git issues, but none that I know of took the plunge. I really think curl should have done that.

      • whatever1 10 hours ago
        Exactly. We invented a pump.

        We can use it as vacuum to clean things up. The same pump can be used as a shit fire hose.

        The effort we need to clean things up is significantly higher than the effort needed to make a mess.

      • hellohello2 12 hours ago
        I think you are correct, its so hard to judge the quality of research that we have been using publication record/count as a proxy. Ultimately, having papers being easy to write is good, provided we find a better way of judging quality.
    • low_tech_love 6 hours ago
      It’s important to notice that flagged papers were still accepted; anyone who spends any time around researchers knows the fact that most people are salivating over AI. That’s why they don’t want to punish it, they’re also using it. And they don’t want to create an environment where it is overly punished, at least not while they can take advantage of it. There is only a small amount of serious people complaining, but in my experience the opinion of the vast majority is “stop worrying and learn to love it”. Which brings to the surface a very harsh reality: almost nobody gave a damn about the science to begin with.
    • strictnein 14 hours ago
      We just need journals run by an AI that charges other AI to read them, and then the AI run colleges can promote the AI with the most AI journal entries and citations.
      • 5555watch 12 hours ago
        I wonder into what weird research niche would all such AI schools converge to after enough time.
        • flir 11 hours ago
          Paperclip research.
      • conception 13 hours ago
        Benchmaxxing is already here. No need to pine for the future!
      • byzantinegene 9 hours ago
        sam and dario would be extremely happy
    • totetsu 11 hours ago
      I have been rolling over the idea of “reasoning deserts” in my head, akin to food deserts. Places where economic calculations lead to only a facsimile of the real thing being provided, without the actual components necessary for human health, wellbeing and flourishing.
    • JumpCrisscross 12 hours ago
      > We are very rapidly automating humans out of the academic publication loop

      We are rendering it irrelevant. If this is the norm for academia, I’m sympathetic to the folks looking to cut its funding.

      • jltsiren 4 hours ago
        This is the norm for the industry. When there is a lot of money to be made, people will chase it by any means they can think of.

        AI is a special case, as academia isn't usually that lucrative. In the rest of CS, many major conferences still have ~300 people, and interesting stuff often happens in specialized meetings with fewer than 100 participants.

        • inigyou 4 hours ago
          This is the norm for the economy. When there is money to be made, at least one person will chase it by any means they can think of.

          Academia is a specific case. It happens everywhere.

      • low_tech_love 6 hours ago
        This is the truth. As a researcher, I say good riddance. It was always kinda stupid, AI is just accelerating the demise of something that hasn’t worked correctly for a few decades now. The time is way overdue for us to figure out some other system.
        • vladms 5 hours ago
          I feel this is a case of "perfect the enemy of good". Looking at the output (scientific progress in biology, medicine, material physics, etc.) the "something" seems to have worked well. Could it be optimized? Probably. Should we completely destroy it and hope a new system will be better? My read of history is that in many cases the new ideas were worse of what they were replacing and it took a long time for a fix.

          So, if you have ideas of a new, better system let's talk about those, before getting happy something gets destroyed and hope someone else will come with a better solution.

          • adastra22 3 hours ago
            Look into the replication crisis. It has seriously impacted some fields very negatively. Psychology and Alzheimer’s research in particular have been set back decades.
            • vladms 2 hours ago
              Look into cancer survival rates : https://www.cancer.org/research/acs-research-news/people-are... , to quote "Decades of cancer research have provided health care professionals with the tools to treat cancer more effectively, so that cancer in general is becoming less of a death sentence and more of a treatable chronic disease,"

              So with the current system, some things work, some things don't work (ex: psychology/Alzheimer). Yes, the system should be improved. No, I am not convinced that "destroying" the current system will result very easily into something better.

              We need to discuss actual solutions for the replication crisis. There are even steps towards improving that, like requiring open data for papers, which makes it harder for people to do some of the manipulations that resulted in the replication crisis. I personally would go even further: you should provide complete documentation (tools, notes, data, raw files, etc.), but then there are some people opposing that due to "privacy" (for medical) or "patents" (for industrial stuff).

              • adastra22 2 hours ago
                What you are implicitly assuming is that the publication and review system matters for that cancer research. Cancer progress, as I understand, is mostly driven by NIH priorities, which are set by governmental review committees. Published work plays a part in that, but it is not the decentralized-review that is typical of other areas of research (e.g. psychology and Alzheimer's).

                The peer review system as we know it today has only really existed for less than a century. It is no how science was traditionally done. It was adopted due to some real problems with the old system, so I'm not saying we should go back. But it's not clear either that unpaid peer-review is the ultimate end-state either.

                100% agreement on your last paragraph though.

      • trunch 10 hours ago
        And make even less knowledge and progress in science public and ever more in the hands of capital and private interests?
        • JumpCrisscross 9 hours ago
          > make even less knowledge and progress in science public and ever more

          If a discipline is spending public dollars at OpenAI and Anthropic, we're funding them with extra steps. (And losing nothing somebody else couldn't do.)

        • a34729t 9 hours ago
          You mean token holders
    • bulbar 9 hours ago
      We try to iterate regarding efficiency (saying that as a neutral observation). Not sure if that works out. As often the case, sciences that don't serve as a foundation for real world results (can this plane fly faster now?) are in much more danger.
    • wombatpm 11 hours ago
      It’s going to go back to the old boys club where personal connections between research groups and institutions will matter more.
    • dwaltrip 13 hours ago
      Unless someone has discovered some magic sauce for getting AI to write well, I have immense sympathy for anyone trying to wade through these papers.

      I may generate slop from time to time, but I do my best to keep it to myself.

      • hellohello2 12 hours ago
        My 2 cents: AI has improved writing considerably for non-native english speakers (in particular China). Writing feels more standardized/boring but easier to read overall. I hit fewer papers that are a pain to read. The most problematic aspect I see are semi-bogus claims i.e. sentences that aren't false, but don't quite feel right either. YMMV.
        • nottorp 8 hours ago
          You mean it may have improved translation...
          • com 7 hours ago
            [dead]
        • breezybottom 2 hours ago
          The grammar may have improved, but it's still just AI slop. I see these garbage papers all the time.
    • plastic-enjoyer 4 hours ago
      > We are very rapidly automating humans out of the academic publication loop here.

      It more looks like the breakdown of the current academic publication system, which was rotten to the core pre-AI and which internal contradictions are just accelerated by AI to the point of breakdown now.

    • logicallee 1 hour ago
      and don't forget, techniques described in papers are now implemented by AI's. For now, a human might point an AI at a paper and ask it to implement and benchmark the technique described there, but the human probably isn't writing the code anymore.
    • porridgeraisin 5 hours ago
      Most things weren't really reviewed that well as it was. It was always a small fixed set of people that actually played their part in the process in good faith. Those people continue to do to today. As a percentage, they were always a small minority. Just that today, they have become even more of a minority.
    • lazide 13 hours ago
      What’s the point?
    • SecretDreams 10 hours ago
      > We are very rapidly automating humans out of the academic publication loop here

      It's a sad thing. But the monetization and enshitification of journal publications over the decade or so, even prior to AI, certainly has not helped this trend.

  • skyberrys 9 minutes ago
    I didn't read carefully enough at the beginning of the article, and at first I thought Caleb and Issac must be two different algorithms for detecting AI in published papers. Eventually I realized they are the names of the two authors who wrote this article (with AI assistance). Anyways, I was wondering how the two of them managed to disagree? Were the papers each human flagged as concerning the same between the two humans? I should keep reading carefully, just thought I would point it out to other humans too, maybe save them from the same mistaken thought.
  • bradley13 5 hours ago
    The next logical step in "publish or perish". Get rid of p-o-p and this problem will also largely disappear.

    I know managers and administrators want a simple metric, but there is no simple metric for research.

    • Viliam1234 1 minute ago
      "Publish or perish" in science is like the "lines of code produced" in software development.

      Even with the technical debt. We have zillions of papers published, we know that most of them probably won't replicate, we don't know which ones.

    • GuB-42 2 hours ago
      But is there a not-simple metric?

      The advantage of "publish or perish" is that it is based on something concrete. The system is gamed, for sure, and even more so with AI, but still, a paper which is cited a lot tend to be useful and the authors get rewarded.

      Remove that metric, and for the lack of a better idea, it will just turn into a game of who has the best connections or who talks the most convincingly, I mean, even more than it is now.

  • low_tech_love 6 hours ago
    “…because conferences have made it mandatory to review 4-5 papers if you submit to them.”

    Is that really a thing? So if I submit a legit paper it is being reviewed forcefully by a bunch of random people from god knows where?

    • emil-lp 5 hours ago
      That's correct. And 4–5 is low-balling.

      I was forced to review 7 papers.

      Now, how are these papers chosen for you?

      You get to bid on which to review.

      Bid on 30 papers out of 30,000.

      I got none of those, and I had to review 7 papers I wasn't really competent enough to review.

      • tgv 3 hours ago
        30k papers? What field is that? How do you even select your area of competence from such a haystack?

        I went to smaller conferences, I suppose, but back then I also had to review papers outside my direct expertise, although nothing too remote. But I didn't know the literature well, obviously. I was able to weed out the sub-par papers (I think), but it was harder to estimate the merits of those that did make sense, as I only saw 1 or 2 per area. And that's what determines the acceptance, after all. No criticism because the reviewer judged it perfect or because the reviewer didn't know what to look for? Outcome is the same.

    • probably_wrong 5 hours ago
      Speaking for the *ACL conferences, they are not "random people" in the sense that they must have "(a) at least two papers in main ACL events or Findings, plus (b) at least one more paper in the ACL Anthology or a major ML/AI venue" [1]. So they are published authors.

      As for it "being a thing": sadly yes. The number of papers has been steadily increasing over the years and the number of volunteer authors is simply not enough.

      [1] https://aclrollingreview.org/incentives2025

  • DarkUranium 13 hours ago
    I feel like this should be treated as, and have consequences similar to, plagiarism.

    Alas, it's probably just wishful thinking on my part.

    • wavewrangler 12 hours ago
      I thought that it was. I thought journals were starting to implement full bans upwards of a year+ for those who don't honestly disclose AI usage in their work? If I'm not mistaken, arXiv is doing this as well? granted, disclosure is different from overuse, but it seems like a small jump to just go ahead and just ban not checking ones work! ...disclosed or not disclosed. my own personal view on it, is if you can't be bothered to spot hallucinations and other such errors, its not ai that is the issue, it is incompetence and laziness
    • azan_ 12 hours ago
      So wrist slap?
      • emil-lp 5 hours ago
        I don't know if you know what you're talking about, but in my area, this could easily result in losing your job.

        You probably don't want to throw someone in jail for one plagiarized paper, so between jail time and losing your job, I don't know what else you have.

  • apwheele 1 hour ago
    It is in alpha (hoping to do Show HN in a few weeks), but for those interested I am working on an application to do this, https://veruscite-data.com/

    Most of the folks on HN will be more familiar with genAI tools and can just use the skill Caleb and Isaac provided in this blog post. My tool is just likely more token efficient and has a GUI where you can review the extracted bib and edit it more easily.

  • kingstnap 13 hours ago
    https://arxiv.org/stats/monthly_submissions

    They should consider swapping this for a log plot.

    I can imagine in 2027 academia looking like Moltbook.

    • voxelghost 6 hours ago
      Since the log trend seemingly started before the age of LLMs papers - cant we just innocently hope that this reflects increasing popularity of prepublishing on arxivx?
    • turtletontine 4 hours ago
      > I can imagine in 2027 academia looking like Moltbook.

      This is certainly “directionally correct”, but keep in mind this is all highly uneven across fields and subfields. The people generating slop articles are mostly trying to publish big flashy things, and naturally ML research has it much worse than most other fields. There are many topics that are super important and interesting, but niche or obscure enough that no slop authors is trying to publish on them yet. So plenty of topics are still dominated by real earnest researchers doing their best, but they’re niche enough that you wouldn’t know about them unless you study that field.

  • dghlsakjg 13 hours ago
    This is a side effect of academia never taking open accessibility to papers and journals seriously.

    If all these papers were not gatekept by journals, it would be trivially easy to validate at least the existence of cited papers and quotes.

    • podocarp 11 hours ago
      I don't think it's academia not taking it seriously, more like these publishers are just coasting on brand name and trying to milk every cent out of it. Peer review is essentially being a Reddit mod or something, you get nothing out of it but karma or a sticker, and the platform/publisher gets all the benefit.

      Imagine being an author and paying to publish your novel. Also btw the editor is another author but has to proofread your book for free.

      Many academics will publish preprints or their more popular papers on their own websites etc. It's just common sense because to academics they don't get paid a single cent by the publisher (and instead have to pay the publishers instead) and more publicity for them is always better than less.

      • thaumasiotes 10 hours ago
        > Imagine being an author and paying to publish your novel.

        That is... completely normal.

        > Also btw the editor is another author but has to proofread your book for free.

        That would be weirder; if you're paying for your own publication, you don't get an editor at all unless you hire them yourself.

      • ModernMech 2 hours ago
        > you get nothing out of it but karma or a sticker, and the platform/publisher gets all the benefit.

        Peer review means 3 people spend their time to review your work and you don’t have to pay them a cent, so you review others' work for free. The compensation for reviewing is reviews.

        They’re not always great (sometimes they’re downright nasty), but I’ve gotten enough high quality reviews over the years that I’m willing to say the time I’ve spent reviewing others’ work has been fairly compensated.

    • crote 13 hours ago
      Validating existence should already be trivial: virtually all journal already have publicly-available indices which contain at least author information and an abstract. Combine that with DOI citations and you're basically done - even with closed-access journals.

      Checking the content is of course a lot more difficult, but that doesn't magically become trivial with open-access journals: you still need to read and interpret what is being said in the paper and compare it to the claims being made in the citation. Granted, these days you could use AI for a first pass, but it's still going to be incredibly tedious work.

    • WarOnPrivacy 13 hours ago
      > This is a side effect of academia never taking open accessibility to papers and journals seriously.

      From my perspective, it's a clear manifestation of humanity's most pervasive failing - the one that defines every group, eventually.

          No One Anywhere Wants To Clean Their Own House.
      • touisteur 13 hours ago
        And yet the automation we pour trillions in, is the one that will do anything, including everything I find interesting and will never clean my own house.
        • lazide 13 hours ago
          Truly passing the real Turing test.
    • nhinck2 13 hours ago
      It is trivially easy to validate the existence of a cited paper.
      • jsrozner 13 hours ago
        I'm sympathetic to the idea: we should have an open, publicly queryable citation graph. Google scholar could very easily offer this at marginal cost near zero, but they won't.
        • AlotOfReading 10 hours ago
          I'd almost rather Google scholar not offer that, because it might draw attention from the eye of sauron and deliver them to the Google graveyard.
        • Kaethar 2 hours ago
          [dead]
  • zenincognito 9 hours ago
  • delis-thumbs-7e 1 hour ago
    I was wondering why they don’t just feed the papers first to an LLM to spot obvious slop, and I was answered later in the post:

    > Submissions are confidential, bibliographies included, and the audit works by sending pieces of one to a hosted LLM – even though the LLM never writes a word of your review. ECCV 2026’s reviewing policies state that LLMs “are NOT allowed to be used to write reviews or meta-reviews, whether it is run locally or via an API,” and separately bar reviewers from sharing substantial excerpts of a submission with an LLM. WACV’s reviewer guidelines call LLM-generated reviews “highly irresponsible behavior,” sanctionable by desk rejection of the reviewer’s own papers, and their confidentiality rules forbid showing a submission’s material to anyone who is not a reviewer – which a hosted LLM is not. NeurIPS’s LLM policy restricts what reviewers can share with LLM services; its AI-assisted reviewing experiment is the sanctioned route.

    So you can send some LLM generated crap to be published and even if you get caught, there is no consequences. Since your funding is likely connected to how much you publish, even if it’s toilet paper, so this system actually rewards one from spewing out shit papers no-one reads. But if you use LLM to review them, guess what, you will get punished harshly.

    I think if you send in LLM crap with hallucinated citations you should get 5 year ban on even sending anything to that conference or publication. And perhaps we should create a local model -based application that filters out this crap. The one the authors had made is a good start, but surely you don’t need Claude to review a bibliography for errors? Surely Qwen with a SearXNG limited to arxiv etc. can do the job?

  • tolugenius 13 hours ago
    > Both papers were accepted for oral presentations with the condition that they simply fix the hallucinated references.

    I do wonder what truthfully could be on ai verification, if even one paper with such an error is accepted it sets the precedent you hopefully get lucky to not get caught (then again verifying for basic tells isn't the same verifying is this genuinely a worthwhile publication, but that's a separate matter)

  • jsw97 13 hours ago
    "A lot of the content of this blog was initially drafted by an agent of some sort"

    What? I mean who does this. My voice is my voice and it's literally never occurred to me to have an LLM do a first draft. I thought that was college kid stuff.

    • xiaoyu2006 12 hours ago
      I can accept having a human first draft and let LLM proofreading and/or do some polish on language, but not the reversed order.
      • volumes94 11 hours ago
        This is an oversimplification and I'll edit. Caleb and I had a conversation about our reviewing woes and thought it would be fun to do an interview style post, so we had Claude come up with some questions based on our convo. We answered the questions from scratch and had Claude proofread at the end.
    • Kaethar 2 hours ago
      [dead]
  • leikarnes 5 hours ago
    Why dont we require the reference PDFs to be uploaded at the same time as the paper?
    • turtletontine 4 hours ago
      This is foolish for several reasons. It’s an undue burden on authors to make them download and upload potentially a gigabyte of files (or more) from many different sources. It is also likely a copyright violation for many or the sources. This also makes it arbitrarily difficult to cite things like conference talks, which are not published texts.

      There are already standard keys to index publications, like DOIs. Requiring a list of DOIs for citations would make a lot more sense and be somewhat feasible, but still doesn’t prevent errors in the author list given in the draft.

  • sumanthvepa 10 hours ago
    Using Pangram to detect AI slop is a very bad idea. I tried it on my own writing which I knew to be written by me and it marked it as AI generated.
    • tibbar 10 hours ago
      I think Pangram is de facto measuring the default Claude voice. It doesn't fire for me on technical GPT writing.
  • mlmonkey 9 hours ago
    > This all is very annoying from inside the review queue. Peer review is unpaid work that we do

    The genie is out of the bottle. We need to figure out a way to contain it. I think a solution is to use LLMs for peer reviews also; fight fire with fire?

    • tdeck 7 hours ago
      Clearly not, if the reason you care is quality.
  • angry_octet 7 hours ago
    I'm concerned that slop authors (or their agents) will use bib-audit in the loop, and hence have perfect references, thereby denying a clear signal of low quality research.
  • myshapeprotocol 12 hours ago
    [flagged]
  • luciana1u 13 hours ago
    [dead]
  • Der_Einzige 10 hours ago
    For anyone that wants to evade the kind of people who want to figure out if an AI wrote your review or not, we wrote a whole paper (ICLR 2026!) on how to do that!

    https://arxiv.org/abs/2510.15061

    I consider all types of "I liked this output, but don't the moment I learned it was AI generated" to be externalizations of "carbon chauvinism" (https://en.wikipedia.org/wiki/Carbon_chauvinism) and basically bigotry.

    And BTW, the term "meritocracy" was coined in a book that was extremely critical of the idea and which argued that a real meritocracy is actually dystopian. We consider our work "harming meritocracy" to be a good outcome: (https://en.wikipedia.org/wiki/The_Rise_of_the_Meritocracy)

  • jeffmanu 12 hours ago
    This is why Eversaid.co is going to be even more useful as ai generated content explodes.