What happened to TheNumbers.com

(stephenfollows.com)

388 points | by nickthegreek 19 hours ago

41 comments

  • abetusk 18 hours ago
    I think people are missing one of the points of the article. It's not just that agents are hammering the site, it's that there might be lurking vulnerabilities that allow malicious usage, which is why it went down, then came back up with a fraction of the data and a reduced design.

    The article says (speculates?) that malicious users are trying to get privileged access for an edge in prediction market betting. From the article:

    > If you could see The Numbers data before everyone else, every single week, you would have a significant edge over all the other traders - learning the answers slightly ahead of publication would allow you to front-run the trades.

    • squidproquo 10 hours ago
      If it's the case that prediction market traders were trying to front-run the market outcomes, the prediction markets should put the fix on their end by locking down trading before a market closes with some time buffer to prevent this.
      • eru 6 hours ago
        Why should that be necessary?

        You can unilaterally stop trading 'before a market closes with some time buffer to prevent this.' No need for centralised action.

        • brookst 31 minutes ago
          Because markets succeed or fail based on perceived fairness.

          It is in a prediction market’s best interest to not become the place where you go to get fleeced.

          Today they’re the Wild West and run like the early days of darknet markets, but if the

      • Cthulhu_ 4 hours ago
        That's assuming the prediction markets all play fair and by the same rules. If Polymarket were to do that, its competition could choose not to. Since these exchanges are not tied to any one country (or jurisdiction), its users would jump ship. Because the users don't care.

        See also stock market trading, where companies would go to great lengths to shave nanoseconds off of getting information and putting in trades. Now apply this min/maxing to worldwide, unregulated and anonymous.

        This is what tech libertarians / cryptobros want.

      • blurgo 9 hours ago
        [dead]
  • podgietaru 17 hours ago
    "One Reddit theory even suggested it was a deliberate rug pull designed to cripple the free site to push people towards paid products."

    I know this isn't really the point of the article, but I've been thinking about this a lot. I wonder if we're going to see more resources go this way.

    I used to publish little doo-dads as opensource software. Not because it was something that was legitimately ground breaking or anything (it absolutely wasn't) but because I had a problem, and I thought "heh wouldn't it be cool if someone else had a similar problem and could use my resource for it."

    But now I'm really reluctant to give more stuff to the free web. Because the fact that it gets scraped and added to a pile of training data to later be monetized really rubs me the wrong way.

    And I can see why maintainers of sites like this, or other free but incredibly useful resources might start to get irate at that.

    • pluc 34 minutes ago
      Why would you offer data for free if people are going to pay someone else to access it?

      The whole internet is about to experience USA-level greed because people are too dumb not to use AI. Just like Google killed small sites because people were too lazy to look at page 2.

    • II2II 17 hours ago
      > But now I'm really reluctant to give more stuff to the free web. Because the fact that it gets scraped and added to a pile of training data to later be monetized really rubs me the wrong way.

      The thing is, people were scraping and monetizing other people's websites long before the current LLM fad. The difference is the magnitude of the problem.

      > And I can see why maintainers of sites like this, or other free but incredibly useful resources might start to get irate at that.

      Those who object to the scraping fall into several camps, but the biggest complaint I am hearing is that it increases both maintenance costs and time. In other words: it sucks when people are using your work in a manner that you find offensive, but it goes beyond that by doing actual harm.

      • TeMPOraL 4 hours ago
        > The thing is, people were scraping and monetizing other people's websites long before the current LLM fad. The difference is the magnitude of the problem.

        > In other words: it sucks when people are using your work in a manner that you find offensive

        That to me still reeks of Dog in the Manger mentality. If you publish something for the world to use, you should neither care nor even track, much less discriminate by (or suddenly seek compensation for) who is using it.

        • Cthulhu_ 4 hours ago
          Agreed, this is what you agreed to (implicitly or explicitly) when uploading stuff to the internet. In fact, open source embraces this (hence the 'open').

          But people are free to not publish things or post things online with a more restrictive license. Not that a license stops things from being indexed.

          • TeMPOraL 1 hour ago
            > But people are free to not publish things or post things online with a more restrictive license. Not that a license stops things from being indexed.

            Right. In fact, people are also free to publish things with licenses that condition access on compensating the author/publisher, and they have both social and legal backing to enforce it. This is called "proprietary", and it's not a wrong choice - in fact outside of software, it's the default choice.

            The problem is when people publish "free" and "open" as a marketing tactic, where in fact they really want to control and charge for access (whether dollars or karma or credit). That is just plain dishonesty.

        • amiga386 3 hours ago
          This is absolutely fine, until some shitheel decides they're going to continually scrape your website, thousands of requests per second, and knock it offline constantly, so they don't have to cache even a single thing, or write better scrapers. They'll just suck up your bandwidth for it. After all, it's you that's paying, not them.

          Search engines were absolutely fine. They were respectful of the sources they scraped. There's now a bunch of money-obsessed shitheels doing whatever they can, as blatantly and as lazily as they can, because they believe they'll become rich. Every one of these fuckers needs to be dead or in jail before the world becomes whole again.

          If nobody takes action against these fuckers, everything you care about will just nope out of existence, and these fuckers will be all you have left.

          • TeMPOraL 1 hour ago
            Major AI companies aren't indiscriminately scrapping any more than search engines were (which themselves caused some problems in the past too, until that got sorted out). Your gripe is with SEO people and content marketers.

            I'd be more sympathetic if it was really about this, though. However, the majority of the complaints seems to come from people personally offended by the possibility of their content, which they claimed was published for everyone to benefit, ending up as training data, and thus actually benefiting humanity, at scale far beyond the original publication ever would. Which is the Dog in the Manger attitude I point out.

      • joshmarinacci 17 hours ago
        A change in quantity can become a change in quality.
      • antisthenes 16 hours ago
        > The thing is, people were scraping and monetizing other people's websites long before the current LLM fad. The difference is the magnitude of the problem.

        People doing it with a couple of machines and residential proxies versus Anthropic doing it with 2 data centers worth of machines.

        Scale matters.

    • chrka 5 hours ago
      I also recently made an open-source project with 200 GitHub stars private. I never had a problem with others using it as the basis for their own projects. In fact, that happened, and I received credit for it. But LLMs just hoover up everything, process it, and then spit it back out as if it were their own.
    • weitendorf 17 hours ago
      I open source as much software as I can because I want the models to train on it and get better at it!
      • Scoundreller 17 hours ago
        I feel the same about my old blog articles.

        I used to rank pretty well trashing crappy credit cards and encouraging people to switch to better options, then Google decided that 10 results for the card issuers website was better.

        Same shit when I manually wrote proto-gethuman posts on calling telecoms/banks/etc (and also pushed visitors to try an Indy ISP or credit unions), then Google felt it was better to drive users to the telecom’s website that wants you to do anything but call them.

        Please do train on my pre-LLM gold!

        • da02 16 hours ago
          Do you have any active sites or social media with your recommendations?

          (I never publish my recommendations because they seem to complicated for people. Like using 30 GB plan/$10/monthly from T-Mobile for data and then using Tello for $10 plan for voice/text. This would require a 2 esim/sim card phone. I am currently using a Moto G Power 2024 phone from eBay $90/new, which was better than the $200 slightly used Pixel 6a from Swappa. Most people would just save the hassle, get a Galaxy phone with a phone contract.)

          • Scoundreller 14 hours ago
            Nothing active anymore, all converted to static sites.

            If you’re on page 2 of the results, you effectively don’t exist so I stopped bothering/benefiting from display ads.

            Coincidentally, I did try to get deeper into the cellular service side (it’s another high margin and high customer value segment), but I did better on the finance side.

    • ryukoposting 10 hours ago
      > Because the fact that it gets scraped and added to a pile of training data to later be monetized really rubs me the wrong way.

      I suspect the next AGPL (if not normal GPL) will explicitly prohibit training of closed models against the source material. Ideally it will prohibit training of open weight models too. Either share the entire process, or go piss up a rope.

      • danlitt 2 hours ago
        If LLM training is found to be a fair use (looks likely), no license will help.
      • dotancohen 5 hours ago
        My understanding of the GPL is that it currently prohibits the training of LLMs, no need to explicitly mention them. MIT does allow it, however.
        • chrka 4 hours ago
          Exactly. Just like using LibGen is prohibited.
    • vachina 17 hours ago
      > But now I'm really reluctant to give more stuff to the free web.

      I feel the same way too. But guess what, all the code you did not publish gets into the training corpus anyway (when you gave Codex or whatever full read access to your filesystem).

      • bigstrat2003 16 hours ago
        That's only if one is stupid enough to give an LLM access to their computer. Don't do that, it's incredibly irresponsible security wise.
        • dotancohen 5 hours ago
          The vast majority of people are in fact this stupid.

          I recently had trouble connecting to a new AP on my laptop. After a quarter hour of frustration, I connected to the AP of my phone, asked Claude Code what the problem is, and she found the issue in seconds. I didn't allow her to make the actual changes, but she did have read access to everything, and helped me considerably.

          So network-manager gets a new bug report about too-long non-ASCII AP names, I get online, and I don't know maybe Anthropic sneakily learned something from my local python projects. I am one of those vast-majority stupid people.

    • protocolture 6 hours ago
      >But now I'm really reluctant to give more stuff to the free web. Because the fact that it gets scraped and added to a pile of training data to later be monetized really rubs me the wrong way.

      I dont get this.

      You provided something for free to help people, but dont want to do that anymore because it might go into training data and help many many more people?

      So far LLMs have been loan funded donations of loss leading services. They might never actually make their first dollar of profit.

      Meanwhile ISPs the world over have been monetizing access to your content.

      I feel like what you mean is that you want control over attribution.

      • danlitt 2 hours ago
        I think the sheer magnitude of the economics have made the scales fall from a lot of people's eyes. For decades people put stuff on the internet for free on the assumption it was "not worth" anything. It turns out that as soon as that commons can be enclosed, we can marshal hundreds of dollars for every single living human, to pay for this commons to be repackaged. The money is there, and we're happy to spend it, we just won't spend it on you.
        • brookst 23 minutes ago
          So basically the “information is free, encyclopedias are expensive” phenomenon from the pre-internet days? Collection, collation, and distribution have more and different qualitative value than the sum of the individual bits.
      • gapan 1 hour ago
        I think he just means that AI companies are not people.
    • CamperBob2 17 hours ago
      This attitude makes zero sense to me. You benefit from the "training data" just like everyone else does. If you don't, that's a problem with you, not a problem with AI models.

      AI solves exactly the meta-problem you describe: "I had a problem and needed to write a one-off doo-dad utility program to solve it." Now you can do something with your time besides writing pointless one-off doo-dads.

      As for monetizing the training data, (a) it cost hundreds of millions of dollars to generate the weights, so why begrudge the companies that made the investment and did the research necessary to make it happen?; and (b) rest assured, whatever your doo-dad does, an open-weight model like GLM 5.2 can generate it for free using your own hardware.

      So you don't have to pay anyone in that case. Well, except nVidia, I guess. Point granted there.

      • akudha 15 hours ago
        It does make sense. When podgietaru wrote their doo-dad programs, I guess they made their work available for humans to use and they didn't mind giving their work away for free. Some other person might say "I'm giving my work away free, I don't care if humans use it or AI, for any purpose". Both are legitimate choices, each according to their own.

        As for "everyone benefiting from training data" - sure, but the AI companies are not investing Billions of dollars out of the goodness of their hearts, it is towards one and only singular goal of making profits (at some point). People might be sympathetic to these AI companies if they at least behave decently - they take everyone's work (text, software, fiction, music, images, videos...) without paying a penny to anyone. If they take everyone's work for free, they should give away anything that is built on that work also for free. This is before we even get to environment, privacy, hammering sites by not respecting robots.txt etc issues.

        If I stole all veggies from your garden but spent time/money/effort making a meal, I should at minimum share it with you for free, no? Even if I spent my own money making the meal, it was made from stolen raw material...

        • protocolture 6 hours ago
          >As for "everyone benefiting from training data" - sure, but the AI companies are not investing Billions of dollars out of the goodness of their hearts, it is towards one and only singular goal of making profits (at some point).

          Its not guaranteed they will ever get there, every day it seems increasingly likely that without massive government intervention we are just waiting for local open weights models to become popular.

          Meanwhile my ISP gatekeeps the same free content behind a service payment. Their motive isnt free love and world peace, its also profit.

          >If I stole all veggies from your garden but spent time/money/effort making a meal, I should at minimum share it with you for free, no?

          If you cloned all the veggies in my garden you are free to clone them further and fill your belly you owe me nothing.

        • CamperBob2 14 hours ago
          If I stole all veggies from your garden but spent time/money/effort making a meal, I should at minimum share it with you for free, no?

          Yes, and that's exactly my point. We are in violent agreement. It's shared with me, with you, with OP, and with everybody else. We can get a delicious meal for free or we can pay somebody else to serve us a slightly-tastier version.

          Here, disregarding copyright law has fulfilled the very purpose of copyright law: to advance the useful arts and sciences. Copyright law was the best tool we had to accomplish that before, and now we have something better.

          • podgietaru 13 hours ago
            I wouldn’t be happy if those companies used my vegetables to open a Michelin star restaurant, even if I got a free McDouble out of the deal.

            Plant based I guess I don’t know metaphors are hard.

            • CamperBob2 11 hours ago
              But the McDouble isn't an adequate comparison. Yes, the companies who opened Michelin-rated restaurants with the food they stole from your garden are giving you McDonalds'-level food for free and charging for the rest. Meanwhile, some other companies who raided your garden are giving you the plant-based equivalent of Ruth's Chris or Fogo de Chão for free.

              The other thing is, after they stole all that stuff from your garden, it was somehow still there. Your neighbors on Hacker News say that some bandits raided your garden, but you can plainly see that no one has picked any fruit or uprooted any plants, and your security cameras reveal nothing more rapacious than a rabbit or two. You begin to suspect that your neighbors are gaslighting you.

              • foco_tubi 8 hours ago
                The plant metaphor breaks down when you pause, take a breath, put the McDouble down and realize that intellectual property is a completely different ownership concept than physical property.
          • cycomanic 13 hours ago
            But the AI companies are not publishing the models? Moreover they are charging for access to the models.
            • CamperBob2 13 hours ago
              Sure they are. Surf around on HuggingFace and you'll see dozens of open-weight models up to 120B parameters in size from for-profit US companies, freely downloadable (if not freely runnable, alas). Probably hundreds of them, at this point.

              And a 3T parameter model is scheduled to be dropped by the Chinese on Monday.

              They should certainly publish more, and if somebody were to argue that model weights trained by scraping copyrighted data should inherently be accessible to everyone, I'd be 100% in favor of that.

      • podgietaru 17 hours ago
        I don't know how else to explain it. It doesn't feel the same to have my contributions smushed into a linear-algebra machine? For me, the incentive was the idea that my contributions might have directly helped someone. That absolutely does not feel the same when I think "My answer has modified the back propagation of a Machine Learning training run."

        It doesn't have to make sense to you - I just believe that I'm not exactly alone in this thought.

        Now apply this to Art, free stories, writing etc. It feels bad to have your free contributions hoovered up and monetized. It doesn't feel particularly fair or ethical to me. And it'd make me double think before making something free and publicly available.

        • CamperBob2 17 hours ago
          Understood, there's certainly nothing invalid about your point of view here. It's just not one that I personally can come to terms with.

          I've spent a lot of time in your shoes, wasting time on busy-work needed to accomplish a larger goal (and absolutely sharing the results freely, over multiple decades)... and I don't miss that part of it one bit.

      • jtuple 17 hours ago
        > Now you can do something with your time besides writing pointless one-off doo-dads

        Writing software one-offs to scratch an itch was historically one of my most enjoyable past-times. AI trivializing that has been a very real theft of joy in my life.

        Solving the actual problem was never the point, it was just motivation to do geek-out and craft some code.

        AI is rapidly diminishing many interesting hobbies (coding, art, music, writing).

        Having more free time when there's nothing fun nor exciting to do with it isn't really a benefit.

      • danlitt 2 hours ago
        > it cost hundreds of millions of dollars to generate the weights

        It cost hundreds of billions of dollars to generate the training data, they just didn't get paid.

      • dminik 17 hours ago
        This is such a sad worldview. Life without any challenges. Every problem immediately solved without learning anything.
        • CamperBob2 14 hours ago
          What's sad is somebody giving you a godlike tool and watching you mope around muttering about "life without any challenges."
          • yuye 10 hours ago
            It's only considered "godlike" by those satisfied with mediocrity
            • CamperBob2 8 hours ago
              The person I replied to is apparently very content with mediocrity.
          • foco_tubi 11 hours ago
            The godlike tools are paid subscriptions at best and restricted to a handful of corpos by the government at best. The free scraps thrown to the proles are mostly novelty. The other day I asked 5.6 Sol to identify a push prop airplane and it said “this is a helicopter.”
      • yuye 10 hours ago
        >Now you can do something with your time besides writing pointless one-off doo-dads.

        This is such a disgusting statement that goes against the very core of what open-source stands for. If even only one person got some use out of what they made, it is not pointless by definition.

      • mtVessel 17 hours ago
        | Now you can do something with your time besides writing pointless one-off doo-dads.

        ...thus eliminating all the tedious chatting, relaxing and making friends that people were previously forced to do while waiting for elevators.

        • CamperBob2 17 hours ago
          ...thus eliminating all the tedious chatting, relaxing and making friends that people were previously forced to do while waiting for elevators

          That also makes no sense, but I don't know what else I should have expected.

      • qotgalaxy 17 hours ago
        [dead]
  • baud9600 5 hours ago
    One thing I’m wondering is whether we need an open source set of technical patterns and libraries for dealing with this changing traffic mix.

    Millions of small sites and creators don’t have the ability to design their own protections against large scale automated access. If useful content now attracts aggressive crawler traffic, many sites will be too expensive or unreliable to run.

    Is part of the answer a community response? Perhaps a community-maintained toolkit, based on traffic data, that host sites can apply? It could include standard agent identification, rate limiting, traffic classification, access policies, caching, challenge mechanisms, logging, attribution and usage control etc etc.

    In effect, we need much stronger road rules for today’s automated traffic, available as open technical patterns and libraries rather than every site owner having to invent this alone (they won’t).

    • m-i-l 4 hours ago
      Yes, a community effort against the botnets would be great.

      The "good bots" which identify themselves and follow the rules are easy enough to block, so not a problem. It is the "bad bots" which pretend to be real users and hide behind residential proxies, and so are almost impossible to block at the moment, which are the problem.

      Given that there are big companies openly (i.e. on the clearweb, not even darkweb) selling access to these residential proxy botnets of compromised smart TVs[0] and mobile phones and other devices, can we not simply get a database of the these IPs and block access from them? I would venture that almost every single one of those residential users are unaware that they have devices in their homes which have been compromised and are being abused in this way, so if they were to start seeing messages from more and more sites along the lines of "Access to this site has been blocked because unusual traffic has been detected from your computer network. Please check all devices on your network and remove any malware which may be routing this traffic." then maybe we could start addressing the problem at the source.

      [0] https://news.ycombinator.com/item?id=49000864

      • TeMPOraL 1 hour ago
        > The "good bots" which identify themselves and follow the rules are easy enough to block, so not a problem.

        And then they get blocked, which is a problem. In particular, the modern agentic AI tools interacting with web services to fulfill user queries - they are acting as user agents, and they should not be discriminated against.

        So I'd say the first pattern that needs to be broadly adopted is non-discrimination of user agents.

        But of course we've tried that in the past, the whole problem is that non-browser user agents == end-user automation, which is anathema to pretty much every on-line business out there, as money made online is primarily conditioned on users wasting their own lives on interacting with services directly.

    • danlitt 2 hours ago
      I don't think there is a technocratic solution to this problem. Very rich people are using their money to DDoS the internet, for no good reason. We just need to identify these people and fine or imprison them until they stop doing it. Residential proxy providers would be a good start.
    • Cthulhu_ 4 hours ago
      The community response was respecting robots.txt, but since that was more of a gentleman's agreement than a legal requirement, hungry AI scrapers disregarded it.

      Which brings us to the old fashioned flood control mechanisms. That is, the toolkit you propose already exists and has been used for decades in various iterations to protect against various forms of attack (slashdotting, DDOS attacks, overzealous search engines, and now AI scrapers).

      Have you looked into those before? Companies like Cloudflare have been at the forefront of this field for a long time now.

  • primitivesuave 17 hours ago
    A couple years ago, I was running a website which allowed the public to view all the US government handouts to small businesses during the COVID-19 pandemic. It also tracked all the fraudulent loans being prosecuted by the DOJ, and allowed anyone to run structured queries over the public dataset. There was a "donate" button which took in ~$2k in donations over the lifetime of the site, and you could download the entire underlying dataset (around 10 GB uncompressed) for free directly on the site.

    Despite the "download all data" link being prominently placed on the front page, the AI scrapers decided it would be more efficient to download terabytes upon terabytes of raw HTML by paginating through every possible facet on the search endpoint. Even with CloudFront caching results and a fairly efficient backend setup, the monthly bill ended up with around $1k just going toward network ingress/egress, so I shut down the site the following month.

    • deepsun 17 hours ago
      BigQuery has "public datasets", so users can even run complex SQL on it, but it's them who pays for it, not you. You only pay for data storage.
      • primitivesuave 13 hours ago
        Thanks for that tip, this is exactly how I would do this if I had to do it from scratch. Just in case it's useful for anyone else: https://docs.cloud.google.com/bigquery/public-data
      • dpoloncsak 16 hours ago
        The crawlers would have still just hammered their site though, right?
        • primitivesuave 13 hours ago
          Yes, if I wanted to put a nice HTML interface over the query results then I would still end up with the same problem where some combinatoric explosion of query parameters to the `/search` endpoint, most of which are cache misses, leads to many many TBs of network egress.
        • pavel_lishin 16 hours ago
          I think the idea is that they could store the data in BigQuery, and point users of the site there.
          • dpoloncsak 16 hours ago
            Sure, but now you're moving the site from "Free data presented in a pleasant way to view" to a "pay-as-you-go database". Your audience shifts dramatically, and you lose the ability to share the data you're trying to present.
            • deepsun 9 hours ago
              No, it's free for everyone for the data sizes they have:

              Free tier: 10 GB of active storage and 1 TiB of query data processed per month.

            • pavel_lishin 15 hours ago
              True. But it sounds like they already lost that.
          • tekne 13 hours ago
            There is also the deep magic...

            https://github.com/phiresky/sql.js-httpvfs

          • PunchyHamster 16 hours ago
            The crawlers would have still just hammered their site though, right?
            • 40four 15 hours ago
              Why is your comment exactly word for word of another comment just one level above in the comment chain?
              • gorgonian 15 hours ago
                Maybe because they restated what they said instead of addressing to the previous commenter’s point.
                • 40four 7 hours ago
                  The comment I’m speaking of was from a different user, and it’s giving a strong smell that both users are bots unfortunately. Maybe I’m wrong, maybe it’s a coincidence those exact strings of characters were formed in isolation from each other in the same thread. I hope I’m wrong, but it’s a red flag
              • taneq 15 hours ago
                Because the crawlers would still have hammered their site, though.

                (The GP post doesn’t actually meaningfully address the issue being raised. Adding BigQuery or whatever would not change the fact that (a) they already offered a method of getting all of the data in a cost effective way, and (b) the issue was that the crawlers hammered the site hard enough to make it economically unviable.)

      • lokar 15 hours ago
        I think snowflake has similar, you can rent it out or make it free
    • sillysaurusx 9 hours ago
      For what it’s worth, a Hetzner dedicated server has unlimited ingress and egress. It seems like it’s the only provider that does. Egress fees suck.

      You can get a beefy one for about $40/mo on their server auction site.

      Just... don’t miss payments. Ever. Or they’ll delete your server within a week or so.

    • userbinator 9 hours ago
      the AI scrapers decided it would be more efficient to download terabytes upon terabytes of raw HTML by paginating through every possible facet on the search endpoint.

      No, they decided that would be a great way to convince you of the narrative and persuade you to pay for "security" services that further the incumbent browser monopoly.

      They're not "AI scrapers", they're DDoS'ers manufacturing consent.

    • tailscaler2026 17 hours ago
      Building on AWS is a financial time bomb.
      • primitivesuave 13 hours ago
        Completely agree, I've worked at several companies where I saw the AWS bill balloon from five to six figures, usually because of pointless over-provisioning. However, it all became worth it when I saw Jeff Bezos go to space. [1]

        1. https://www.youtube.com/watch?v=IOmX793-5t4

      • ninalanyon 16 hours ago
        Are you not able to put limits on how much the site can spend?
        • mystifyingpoi 15 hours ago
          As of 2026, still not, and probably never.
        • shermantanktop 16 hours ago
          Sure, but many a hobbyist has discovered the need for that the hard way.
    • BeeOnRope 17 hours ago
      Do you have any view on why the AI scrapers resulted in a heavier load than existing crawlers from eg search engines?

      Where they more exhaustive or more frequent?

      • primitivesuave 13 hours ago
        They were both more exhaustive and more frequent. I don't remember the exact numbers, but it was definitely over 100x the traffic from search engines. By the way, search engines were allowed under robots.txt - I did want all 11.5 million loans to be individually indexed so they would pop up in Google search results, and I actually did receive/forward multiple tips about fraudulent loans because someone searched a business name on Google. All of this traffic was barely a blip, and my AWS bill for my hobby data science account was only ~$50-$100/month.

        When I looked at the logs after getting the billing alert, 99.99% of the requests were to the "/search" endpoint with virtually every permutation of ~10-12 facets in the query parameters. There was only one scraper, but it triggered an enormous amount of network egress since it ended up missing the cache on the majority of queries.

      • pverheggen 16 hours ago
        There's been an explosion of vibe-coded scrapers that behave poorly and ignore robots.txt. Presumably OP disallowed crawling of the search endpoint for the reason stated (that crawling it would result in endless permutations of search filters.)
      • ratelimitsteve 17 hours ago
        not the OP but I'd say that google crawls you once and AIs scrape your page every time someone asks them a question that they think your page might be relevant to.
    • x3haloed 15 hours ago
      That’s a traffic design problem. You should be happy that your work is valuable and also protected it against excessive requests. Simple.
      • primitivesuave 13 hours ago
        I deployed a Lambda function behind CloudFront which rendered a simple HTML page with the query results from executing some SQL over the dataset. I served millions of page views for next to nothing because most pages were already in the cache.

        I don't know where you get this expectation that people should anticipate that a crawler might try every possible combination of query parameters, thereby missing the cache on each one. Most people consider it a bitter and arrogant perspective, which is why this got downvoted.

  • breve 7 hours ago
    AI companies really do socialize the costs and privatize the profits.

    Sites like The Numbers have to take on the cost of surviving the AI onslaught and the AI companies return nothing back to them.

    • Cthulhu_ 4 hours ago
      To a point; as the post mentions, they added instructions for LLMs to their website and increased their licensing inquiries tenfold (the article didn't state anything about actual licensing payments being made though).

      I don't think there would be any issues per se if the scrapers just paid licensing fees to get the good / complete data. But the issue was that they started to try and find exploits to get to data earlier.

    • jambalaya8 6 hours ago
      worse; it isn't just no return, it is also cost.
    • daniela-scott 5 hours ago
      [flagged]
  • djoldman 15 hours ago
    Just FYI, the bigger companies all allow you to block crawlers via robots.txt:

        # Block Anthropic (Claude)
        User-agent: ClaudeBot
        Disallow: /
        User-agent: Claude-SearchBot
        Disallow: /
        User-agent: Claude-User
        Disallow: /
    
        # Block OpenAI (ChatGPT)
        User-agent: GPTBot
        Disallow: /
        User-agent: OAI-SearchBot
        Disallow: /
    
        # Block Perplexity
        User-agent: PerplexityBot
        Disallow: /
    
        # Block Google's AI Training
        User-agent: Google-Extended
        Disallow: /
        User-agent: Google-Extended-Factual
        Disallow: /
    
        # Block Microsoft's Search & AI Crawler
        User-agent: Bingbot
        Disallow: /
    • Cthulhu_ 4 hours ago
      This only works for crawlers that play fair, unfortunately there's many that don't. Including ones that use consumer devices like TVs [0] and mobile phone apps that offer incentives to consumers (if they're even open about it) to use their internet connection to e.g. crawl websites. There's probably an army of hacked toothbrushes and the like doing the same thing.

      [0] https://news.ycombinator.com/item?id=49000864

    • frereubu 15 hours ago
      This is good to know, but a bit of all-or-nothing. It's a shame that, for example, Google doesn't support the crawl-delay field so you can tailor their crawling to your setup: https://developers.google.com/crawling/docs/robots-txt/robot...

      I presume it would also cut you off even more from referral traffic.

    • pixelesque 13 hours ago
      Claude Bot still (at least last month, and it's been doing it for over a year now) seems to have a bug when traversing (at least my sites), wherein it drops the trailing slash of a directory (which is present in the a href tag), then makes the request to the subdirectory without the slash I put in the link, then Caddy automatically responds to that (via the built-in File handler) with a redirect telling it to add the trailing slash, and Claude Bot then makes the request again with the trailing slash that should have been there in the first place.

      So most subdirectory URLs get two requests from Claude bot, the first one needless because that wasn't the URL in the tag.

    • spiderfarmer 14 hours ago
      90% of bot traffic on my network of websites is through headless Chrome, via residential bots nowadays. Impossible to block. Not even for Google, as they inflate my Adsense numbers as well.
      • 6510 14 hours ago
        I once had the dumb idea to make websites in pdf but the idea is growing on me. Perhaps it should even be in animated gif with <img> <map>'s
    • globular-toast 5 hours ago
      I won't eat your lunch if you put a sticker on your lunch box telling me not to.
  • ethagnawl 18 hours ago
    At the risk of oversimplifying things from a distance, this site -- especially the free, public-facing part of it -- seems like it would be an ideal candidate for a rewrite using static site generator/framework. That, coupled with a bot-aware CDN should keep them online at a reasonable cost for many years to come.

    Otherwise, I'm very curious to know more about their old and new architecture and what sorts of mitigation/scaling strategies they've started using to keep the site online.

    • zackmorris 17 hours ago
      Ya this problem was solved 20 years ago with Varnish cache and Coral cache/CDN.

      I think what's really going on is that bots expose how underpowered web servers has gotten in recent years. In the 2000s, even poorly-architected PHP sites tended to serve about 200 requests per second, with 1000+ being common for static sites. I remember when Node.js came out and claimed that it could serve more like 100,000 RPS due to its cooperative threading model. But today sites have a remarkable slowness to them, running many hundreds or thousands of database queries due to ORMs and N+1 problems, so that response times can be 500 ms or more and even 1000 simultaneous users stresses servers.

      What went wrong is that nobody solved stuff like Russian doll caching in a general way separate from the programming language and database. We should have had ways to make dependency graphs using Etag headers as keys with real cache invalidation of dependent data. I went down that rabbit hole 10 years ago using touch events in Laravel with callbacks to handle cache invalidation when class model data was saved to the database. Also a query cache using Redis which I think might have been handled better at the database level anyway. After that experience, I can honestly say that cache invalidation is so difficult to get right that it's effectively an open problem. Meaning that programmings should use a package instead of rolling it by hand, and it should be a major concern from the start (along with sharding by user id or using something like Firebase).

      Don't get me started on how the web should have been a P2P content-addressable memory anyway. Nearly everything should be available from a nearby edge peer, similarly to BitTorrent. But nobody bothered to solve how to make that work with HTTPS/SSL. I suspect that has to do with early flaws in the browser security model where the whole page has to be behind HTTPS or warnings appear. So it was never clear what was personally identifiable information (PII) or merely public data being served over HTTPS. To really solve that, we probably need real trust networks and maybe even zero-knowledge proofs.

      Since these problems are so challenging to fix, and big companies can't be bothered to do it since they pulled the ladder up behind them, we're probably stuck with banal "are you human" challenge screens for the foreseeable future.

      • withinboredom 17 hours ago
        They are indeed quite challenging. We literally wrote a paper about it (https://arxiv.org/abs/2605.09114) ... you almost describe some of our original architecture too! You might find it interesting -- or not. https://getswytch.com
      • atherton94027 11 hours ago
        > I think what's really going on is that bots expose how underpowered web servers has gotten in recent years

        More often than not they're running on a VPS, and cloud providers have been pushing the envelope of what a vCPU is. Amazon still bills hyperthreads as a single core!

        Add to that their servers are often aging and you get a recipe for slow web services

      • Terr_ 15 hours ago
        > Nearly everything should be available from a nearby edge peer, similarly to BitTorrent.

        Back in the days of Gnutella, I remember pushing people to use Magnet links [0] when sharing content.

        [0] https://en.wikipedia.org/wiki/Magnet_URI_scheme

      • PunchyHamster 16 hours ago
        > Ya this problem was solved 20 years ago with Varnish cache and Coral cache/CDN.

        It was... if you are paying datacenter rates for the bandwidth

        If you're paying cloud provider per GB pricing, nah, even text will add up if you happen to be targeted by a bunch of bots

        > What went wrong is that nobody solved stuff like Russian doll caching in a general way separate from the programming language and database. We should have had ways to make dependency graphs using Etag headers as keys with real cache invalidation of dependent data.

        we did that with nested ESI includes in Varnish so every "box" of content on the page was cached separately + some piping for invalidation, so if a given piece in the database was changed it sent invalidation to all nodes. There was also some grace so if the thing you wanted got updated RIGHT NOW you might get stale version while the new one is updated in background, and don't pay the latency cost

    • tengada1 18 hours ago
      Yeah this seems like one of the easiest access models to create an excellent security model for -- basically just static publishing. Sounds like they were in need of a rewrite anyway!

      It probably seems daunting but to be honest this feels like a weekend's work at this point with LLM assistance. Not to be glib!

      • jaredwiener 18 hours ago
        But it also destroys the business model behind the site.

        Sure, you could technically redesign to handle the bot traffic, but if the bot traffic is just taking the data and reducing any need for humans to visit the site, why is he putting in the effort to maintain the site?

        • tptacek 16 hours ago
          The business model of the site is apparently private data sales, not ad revenue.
          • jaredwiener 15 hours ago
            Bots don't make purchasing decisions.
            • FinnKuhn 1 hour ago
              According to the post inquiries for purchasing the data increased as a result though so they might not make the decision, but they do make the suggestion.
            • Terr_ 15 hours ago
              Some people are already trying to make that happen, but I'm convinced it'll be a net-negative.
              • jaredwiener 15 hours ago
                Rephrase: Bots SHOULDNT make purchasing decisions.
                • fragmede 9 hours ago
                  Because... ? If I'm operating a site, and I want $X to allow a bot to scrape my site, why shouldn't the bot be allowed to make the purchasing decision to scrape the site? Obviously the bot owner would have their own set of guardrails, but if it allows sites like The numbers to stay up because bots aren't going to look at ads so the old model of displaying ads isn't going to work, I have a hard time seeing that as a bad thing.
                  • jaredwiener 8 hours ago
                    I mean it from the consumer perspective.
        • hyperhello 18 hours ago
          Well, why is he? If it's for ad revenue, this problem is surely global to the web. If it's because he likes to do it, it shouldn't matter if one AI or a trillion scrape the site. I feel like we badly need a rethink of the web architecture after DoubleClick anyway. Maybe the name of the game should be to cut down a site's assets very hard and use static hosting for them. This interferes with crummy sites that show a mess of inlined ads every refresh, but now that there are a billion 'poor users' this is no longer feasible.
          • jaredwiener 17 hours ago
            Maybe more an inherent problem with these chatbots.

            LLMs are great, but they aren't producing new information. You still need people for that. But if you cut down any incentive for the people to do that, the LLMs will starve.

    • RobotToaster 17 hours ago
      Or just use a cache.
  • theragra 14 hours ago
    I have small website with archive of older radio broadcasts mostly in Russian. Lately, 90% of traffic comes from USA ;)

    Previously it was like 5%.

    URL is http://radar.lv btw

    Thanks God i can serve up to a terabyte per month of traffic easily, otherwise it would be a catastrophe. I am not against bot scraping, but I worry about stability for meat visitors.

    So, I think of enabling payments for website visits cloudflare recently developed.

    I also have the problem of old technology like the mentioned site. While my personal blog uses static generator, archive website uses ancient Drupal version, which has no security patches for many years already.

  • gajus 18 hours ago
    What a throwback. Back in 2015 I have started Applaudience, which at the time was the only provider of real-time cinema ticket sales data. I still remember comparing our numbers against TheNumbers.com as part of calibration. I have since moved on to other businesses, but this remains one of my favorite pieces of technology that I have developed. Would love to bring it back one day.
  • rodarmor 9 hours ago
    There’s an easy solution: Bruce should place large bets on the prediction markets right before publishing the relevant data. It is legal, ethical, and, for him, risk free.
  • ajkjk 16 hours ago
    It feels like there is fundamentally missing infrastructure here that is needed to make these problems go away.

    Basically bots need to be (somehow) paying for the traffic they create, or prevented from creating it, or told to go away and then fined if they violate the request. No idea how to do these or even at what level in the stack they should happen, but they need to happen eventually somehow.

    • polio 16 hours ago
      Every website I visit could get a fraction of a cent in my Cloudflare Wallet. A human browsing incurs a few dollars a month. Plus, any website I access in this way decides not to show me ads either and just charges me the cost of serving the data plus a nominal profit margin.

      Disclosure: I am a shareholder and would love for them to solve the AI bot problem and the ad problem like this.

      • wredcoll 11 hours ago
        Maybe we'll all eat our words and cryptocurrencies will actually become useful for something.
    • john_strinlai 16 hours ago
      we've been working on basically the same problem in the email space for a few decades (legit email vs. mass spam). its a very hard (i think impossible) problem.
      • ajkjk 8 hours ago
        Well, it is technologically not that hard to imagine a solution---the problem is social, getting everyone to agree on how to do it, figure out the policy around it, etc. And the situation is getting so bad that maybe it is time for someone (maybe someone reading this thread) to figure it out.
      • yuye 9 hours ago
        The same goes for telephony.

        I reject all phone calls by default, unless I'm expecting a call.

  • PowerElectronix 4 hours ago
    Only solution I see is to ban user that abuse the system. Too much requests, too much bandwidth and you get into the blacklist and are cutoff from the human side of the internet.

    With the amount of money AI labs are burning, someone should just set up some infra to host the data they want and charge for access to it, instead of abusing the goodwill of legitimate pages that offer it for free.

  • paxys 17 hours ago
    I know people have opinions about Cloudflare but why not use it here, at least as a stop gap? Stopping bot traffic is one thing it does very well.
    • philipkglass 13 hours ago
      I set Cloudflare up a couple of months ago specifically to block bot traffic. It didn't do anything for me. Dumb bots were still hitting every special link on my wiki fast enough that the server was continually swamped running Lua scripts. 65% of the traffic for my English-language site was coming from Vietnam. But I didn't want to block Vietnam altogether, because my hobby site has genuine users from there too.

      I eventually just shut down my mediawiki instance. I couldn't find a way to keep it online and still run on an affordable VPS.

    • spiderfarmer 14 hours ago
      It tried it. Hundreds of “genuine” visitors per day on a new website with no search engine presence and no links. That’s a very leaky fence..
      • paxys 14 hours ago
        Hundreds is better than billions.
    • NetMageSCW 17 hours ago
      Cost?
      • esseph 17 hours ago
        It's free for this
  • cogogo 16 hours ago
    I think I probably use AI like a lot of consumers out there. Search engines have gotten bad and AI really good at answering fairly specific questions. Often pointing at sites like wikipedia. Pretty clear changing my behavior will have zero impact but it is certainly part of the problem. Feels a lot like my CO2 consumption.

    *Edit - CO2 creation

    • spiderfarmer 14 hours ago
      GPTBot crawled millions of pages on my website. I get a handful of visitors from them. Meanwhile Google sends me 10k visitors each day.

      So GPTBot is now blocked.

  • drumhead 2 hours ago
    The Numbers is such a brilliant site, this explains why it came back in such a stripped down form. AI and prediction markets giving me more reason to hate them.
  • Centigonal 8 hours ago
    Our internet ecosystem is becoming more and more hostile to open information commons. I like open information commons -- what do we do about this?
  • datadrivenangel 18 hours ago
    If the AI companies destroy the open web, eventually they'll need to start curating knowledge sources just like netflix makes movies and amazon has physical stores...
    • nradov 18 hours ago
      Already happening. Frontier LLM vendors have been hiring human domain experts specifically to create their own proprietary training data in targeted verticals.
  • Drdiamond 7 hours ago
    This is such a sad worldview. Life without any challenges and solved using AI. it will vastly decrease our learning
  • xeyownt 14 hours ago
    Prediction market as well as anything related to gambling must disappear from earth surface. Full stop.
  • tehjoker 17 hours ago
    The real story here is that prediction markets were banned for a reason and loosening the rules is causing chaos just as was expected. AI plays little role in this story, hacking by humans would also be motivated by financial returns, unless the element is that AI hacking is cheaper and the returns are not so big.
  • tailrecursion 15 hours ago
    Do the AI labs in the U.S. publish the AWS IP address sets that their crawlers use? Is that not an effective way to block those crawlers anymore?
    • theragra 14 hours ago
      They are often from residential IP.there are even businesses that allow you renting such IPs
    • weird-eye-issue 11 hours ago
      Nobody serious would ever use an AWS IP lol
  • hebleb 17 hours ago
    Good read, I was so curious how this happened a few months ago
  • advisedwang 17 hours ago
    This article, and possibly the owners of the site, seem to muddle together several problems:

    1. bot traffic causing infrastructure cost

    2. scraping circumventing paying for licenses

    3. the risk of hacking

    • shermantanktop 16 hours ago
      You missed the speculated motivation: unrestricted prediction markets, which are an open and broad incentive to do whatever actions might provide a slight edge in betting.

      Webscraping is among the more benign things that gambling-on-anything can drive. And even that has a negative impact, as seen here.

      I have no idea why those sites are legal.

  • jagermo 4 hours ago
    prediction markets were a huge mistake.
  • dumberquestions 18 hours ago
    Ironic that polymarkets were being advertised as helping society make better predictions.
    • jambalaya8 17 hours ago
      John Brunner and Alvin Toffler both made stark warnings wrapped in futuristic giddiness about things like it. People like Fuller no doubt thought polling on a large scale was terrific. There were old usenet groups and BBS subs (minus the money aspect) experimenting with the model. I do not believe they ever are or were good in a largescale model (money or not).
    • mrandish 17 hours ago
      I mean wisdom of crowds, super-forecasters, calibration and pre-registration are useful tools that can result in better predictions, turning it into online gambling is where it went sideways.
      • Terr_ 15 hours ago
        Mini-rant: Promoters claim that the system serves a public good by helping society discover/converge on useful truths sooner.

        Even where the question/event is of public interest, the opposite happens instead. People with expertise or non-public information are incentivized to misdirect and delay as long as possible, as that maximizes what they can make from betting.

  • hyperhello 18 hours ago
    If the scrapers are going to get it anyway, put the data up as a zip somewhere.
    • jjgreen 18 hours ago
      It would not make any difference.
      • codemonkey-zeta 18 hours ago
        Indeed, the article mentions Wikipedia experiencing similar scraping pains, even though they already DO have bulk data available.
        • HeatrayEnjoyer 17 hours ago
          Who are running these bots? I presume developers at all of the frontier labs know (or at least would know to look for) Wikipedia has bulk APIs for automated access. Unnecessary scraping increases their workload/costs too, so why in 2026 is this still a problem?
          • esseph 17 hours ago
            Black market and gray market data. All the firms want data. All the other firms want data. The banks want data. The other criminals also want data for their crimes and schemes. Oh insurance companies, and the ATS systems. Everybody wants as much data as they can get and they don't care how they get it.
            • esseph 15 hours ago
              This data selling also happens with leaks of all kinds like medical data, often to current or future employers, health insurance companies, etc.
    • antisthenes 16 hours ago
      > Read the Docs, a non-profit that hosts documentation for open-source software, who watched a single crawler download 73 terabytes of zipped HTML in one month, costing it over $5,000 in bandwidth

      From the article.

      Not the same site, but an example of the same issue.

  • freediddy 15 hours ago
    All new content will be behind paywalls. This will be the new way the Internet works and I've said this for the last several years. There's no value in writing unique content just to have AI steal it and distribute it without you getting any clicks. The only way it works if it there's a licensing agreement and if they don't want to pay, then who cares, you weren't going to make any money from them anyway.
    • npilk 15 hours ago
      What if you aren't motivated solely by making money? There will still be new free content created by hobbyists, passionate creators, altruists, etc.

      (Agree with you more generally.)

      • freediddy 15 hours ago
        Why would you produce content and have no one read it or visits your site, but OpenAI and Anthropic make millions from it? At some point it becomes stupid to just give free money to these companies when they steal literally all your traffic and content. As per the article, Anthropic sends 1 page view for every 38,000 views they get.
        • ssl-3 13 hours ago
          Very few people showed up to read stuff on the old web, and yet: It existed.
  • puskavi 10 hours ago
    Eh, wouldnt just adding some POW challenge, like anubis solve the scraping by making it hurt crawlers wallets?
  • jambalaya8 18 hours ago
    I always liked this site, but reading this and seeing the anger about expecting the site maintainer to do things for you is repulsive. Frankly, if he wanted to pull his site down with no notice that is perfectly within his right. It was/is his site. He doesn't owe anyone a .tar.gz either. His work.
    • BraveOPotato 18 hours ago
      I agree. I've seen it happen before on a project I use. I decided to take a look at the repo for one of the plugins, and I saw a heinous issue that basically was TELLING (not even asking) the maintainer to fix it.

      Deplorable behavior indeed

      • jambalaya8 17 hours ago
        Like dominoes, as soon as it is accepted in a few places, people think it is acceptable to push to 'share'. It's almost terroristic sometimes, the pressure some maintainers are under.
  • ropable 11 hours ago
    This is another reason we can't have nice things. The level of entitlement required to send an angry message to someone complaining about their free resource being offline is mind-boggling, though.
  • NetMageSCW 17 hours ago
    What does a cyber attack have to do with AI scraping?
    • john_strinlai 17 hours ago
      they both have a risk of harm which the site operator was no longer comfortable with.
    • Good4boothee 2 hours ago
      [dead]
  • BrenBarn 8 hours ago
    Every example of this confirms my view that we should treat the spread of these AI scrapers like we treat the proliferation of drugs. We should be seeking to bust AI rings like we seek to bust drug cartels.
  • brcmthrowaway 18 hours ago
    Its just too easy for a technically minded bored person to produce slop that hammers websites. There needs to be a penalty for this.
  • cawksuwcka 11 hours ago
    [dead]
  • draw_down 19 hours ago
    Sorry! Just yesterday many of us decided that AI scraping isn't a real problem, and anytime it is blamed, it's a cover for something else.

    https://news.ycombinator.com/item?id=49005747

  • aaron695 17 hours ago
    [dead]
  • VulgarExigency 16 hours ago
    They should sue Anthropic for this distillation attack
  • vachina 17 hours ago
    > The world we have built thus far is so incredibly ill-prepared for the power and scale of the AI models we all have access to.

    Exactly, so use the AI to secure your servers. Ask the AI to audit your site for any security holes. If unable to rewrite, at least harden the existing code. AI’s are really really cheap (and fast) security consultants now.

  • nater5000 18 hours ago
    This is an odd example to present this argument through.

    I'm not defending AI bots overwhelming websites or hackers motivated by Polymarket, etc., but I don't really believe a 30 year old website with "approximately 160,000 source files serving around 2 million pages" is a good litmus test for the state of the online world. What's worse is the absurdity that this basic site offering niche data would be targeted because Polymarket depends on it for some of their bets, something that most websites don't have to deal with. Frankly, that seems like a much more interesting angle to explore than "this old website that should have been re-written multiple times over the last three decades now doesn't have a choice but to be re-written."

    Oddly, it seems like the solution, at least in this specific case, is relatively simple (and ironic): pay for an LLM to re-write this site in a modern, scalable, secure way and have that LLM monitor the site to ensure it is behaving correctly. Put this thing in a modern platform behind a proper cache (i.e., throw the whole thing up in Cloudflare) and many of these problems just don't exist anymore. If this site is as basic as it seems, this could be a weekend project.

    It's certainly an interesting story, and I don't blame the owner of The Numbers for handling his site the way he has, but framing this as "AI is destroying our beloved internet" just seems obtuse, at least through this lens.

    • jambalaya8 18 hours ago
      The actual solution is damn simple and painfully obvious: make things like so-called "prediction markets" illegal.

      The intelligence community came to the conclusion (after research and experiments) that things like "prediction markets" were a bad idea in 1996.

      Far worse now than then, with the net and AI of 2026 and the insane number of people now online that want to make an easy buck.

      Don't get me wrong; I am guessing some of us on here would do well on those places. But they should not exist.

    • jaredwiener 18 hours ago
      Isn't this the NRA's argument? The only way to stop a bad guy with a gun is a good guy with a gun.

      Or at least somewhere between that and full on protection racket.

      LLMs come on the scene, hammer the site until it goes offline or racks up bills that threaten bankruptcy, and the solution then is to pay for an LLM to fix it and monitor it.

      That's a real nice site ya got there, it would be a shame if massive datacenters going up around the world were to start hammering it from tens of thousands of IP addresses....