Tuesday, August 25, 2026

Chatbot Usage Limits Now are Effectively Unlimited for Most Models for Light Users

Generative artificial intelligence model usage allowances for users on free plans have changed substantially since the first generation of each major model, generally following the pattern for internet access services: moving from usage limits to effectively unlimited for light users. 


ChatGPT might not have had formal usage limits, but access was the real constraint: the model was at capacity so often that many users found they could ask a few questions before hitting an effective block. 


Platform & Model at Launch

Initial Free Tier Launch Date

Initial Usage Limits (When First Introduced)

Reset Window / Conditions

OpenAI ChatGPT (Original GPT-3.5)

November 2022

No hard rigid prompt caps initially, but subject to broad error messages ("ChatGPT is at capacity right now") when servers were overloaded. Users could typically send dozens to hundreds of messages freely.

Dynamic system load throttling; no rolling time window concept at day one, just global traffic blocks.

Anthropic Claude (Original Claude 1)

March 2023

Roughly 50 to 100 messages per day depending on server traffic, as Anthropic quietly tested its early constitutional AI assistant against a smaller user base.

Reset daily at midnight.

Google Gemini (Originally launched as Bard using PaLM 2)

March 2023

No explicit hard numerical cap on prompt quantity for the web app interface during its initial experimental rollout, though safety filters and length restrictions applied.

Controlled dynamically by global capacity limits rather than a rigid per-hour user meter.

Microsoft Copilot (Originally launched as Bing Chat using GPT-4)

February 2023

Initially capped strictly at 5 turns per conversation and 50 queries per day to prevent erratic behavior and control high compute costs of early GPT-4. (Limits were quickly relaxed to 20/300 after initial tests).

Reset daily; individual chat sessions had to be wiped clean after hitting the turn limit.


Today, as more capacity has been added, some operations, such as text queries, are unlimited in principle on ChatGPT free plans, for example. 


Lighter users might seldom, if ever, encounter access blocking because servers are at capacity. 



Platform

Model Access (Free Tier)

Message / Usage Limits

Reset Window / Conditions

OpenAI ChatGPT

GPT-5.6 Luna (or equivalent lightweight default model)

Unlimited text chats (introduced for text as of August 2026); limits apply to heavy features like file/image uploads, voice mode, and image generation.

Weekly rolling limits or dynamic caps apply to resource-heavy features (like advanced tool usage or deep searches).

Anthropic Claude

Sonnet-class model (e.g., Sonnet 4.5)

Roughly 15 to 40 messages per rolling window. Limits are token-dependent (long documents or code pastes consume the quota much faster).

Rolling 5-hour window (refills continuously 5 hours after your first message; no midnight reset).

Google Gemini

Gemini Flash / Flash-Lite variants

Generous message allowances for general prompting; API free tier allows 5 to 15 Requests Per Minute (RPM) and up to 1,000 Requests Per Day (RPD) depending on the specific model.

Standard rolling limits apply on the consumer web app; API limits reset daily/continuously.

Microsoft Copilot

GPT-4o / latest integrated OpenAI flagship infrastructure

Standard daily chat caps (typically ranging around 30 to 50 turns per conversation, with a total daily cap around 300 messages depending on demand).

Resets daily or after starting a fresh conversation session.

Sunday, August 23, 2026

Resistance to New Data Centers Might be a Developing "Moral Panic"

Moral panics sometimes occur when a society has anxieties about modernization, shifting social roles or perhaps new technologies. 


Opposition to the building of high-performance data centers supporting artificial intelligence operations provides a possible case in point. 


Yes, there are land use, electricity and water consumption issues. But the actual effects or impacts often are exaggerated.


Data-center electricity use is rising fast because of AI, to be sure.


Globally it was about 1.5% of electrical demand in 2024 (about 415 TWh) and is projected to roughly double to three percent or so by 2030 under IEA base cases, with AI-focused facilities grow faster. 


And data center power demand represents a large fraction of incremental demand growth.


Electricity growth is still a minority of overall demand growth globally and competes with air conditioning, electric vehicle charging, industry, and electrification in general. 


Category

Approximate Annual Electricity Use (US)

Notes / Equivalence

Data centers

~176 TWh (2023); ~180–192 TWh (2024)

~4.4–4.7% of total US electricity. Rising rapidly with AI; projections for late 2020s/2030 often in the 300–800+ TWh range depending on scenario (9–15%+ of US total in higher cases).

Residential / Households

~1,400–1,500 TWh (order of magnitude)

Average US household ~10,000–11,000+ kWh/year. Data centers currently comparable to the electricity use of roughly 15–18 million average households.


Water use for cooling is a big issue in drought-stressed or aquifer-dependent areas. And noise, land conversion, diesel backup generators, and rate impacts (when grid upgrades or generation are socialized) add to the list of issues.


On the other hand, U.S. data-center water use remains a small fraction of total freshwater (perhaps half a percent) and is dwarfed by water used for golf courses and agriculture.


Category

Approximate Annual Water Use (US)

Notes / Equivalence

Data centers (direct consumption)

17–17.4 billion gallons (2023)

Lawrence Berkeley National Laboratory and related analyses. Projected to rise (e.g., 38–73 billion gallons by 2028 in some scenarios). Mostly evaporative cooling. Nationally <0.5% of freshwater use. Equivalent to roughly 160,000–580,000 households (varies by per-household assumption).

Golf courses

~531 billion gallons (2024)

1.63 million acre-feet (GCSAA survey). Down ~31% since 2005 due to efficiency and fewer courses. Roughly 30× data-center direct use.

Agriculture / Irrigation

~26.4 trillion gallons (2023)

81.0 million acre-feet applied (USDA NASS 2023 Irrigation and Water Management Survey). By far the largest category; irrigation accounts for roughly 40–50% of total US freshwater withdrawals in recent USGS data. Orders of magnitude larger than data centers.

Residential / Households

Roughly 10–15 trillion gallons (order of magnitude)

Public-supply domestic deliveries are a major share of the ~35–39 billion gallons/day public-supply total. Average household often cited around 100,000–120,000 gallons/year (varies widely by region, family size, outdoor use). Data centers equal a very small fraction of total residential use.


So some might view the concerns as legitimate, but wildly overblown, especially considering all the other value AI and high-performance computing might represent across the whole economy in reducing resource impact. 


Water use for cooling is a big issue in drought-stressed or aquifer-dependent areas. And noise, land conversion, diesel backup generators, and rate impacts (when grid upgrades or generation are socialized) add to the list of issues.


On the other hand, U.S. data-center water use remains a small fraction of total freshwater (perhaps half a percent) and is dwarfed by water used for golf courses and agriculture.


And such concerns do not include benefits such as reduced resource consumption in all other areas of the economy affected by AI and high-performance computing.


Sector / Domain

AI / HPC Application

Mechanism of Resource Reduction

Illustrative Potential Impacts

Agriculture

Precision irrigation, nutrient management, crop monitoring (sensors + satellite + ML models)

Apply water, fertilizer, and pesticides only where and when needed; detect stress early

Water savings commonly 20–50%; fertilizer reductions ~25–30%; higher yields on same or less land

Energy systems / Grids

Demand forecasting, renewable integration, predictive maintenance, flexible load management

Better match supply and demand; reduce curtailment of renewables; shift or curtail flexible loads; avoid unnecessary generation and infrastructure

Lower peak demand, higher renewable utilization, reduced need for peaker plants and excess capacity

Manufacturing

Digital twins, process optimization, predictive maintenance, quality control

Simulate and optimize processes before physical runs; minimize scrap, downtime, and energy waste; right-size material and energy inputs

Energy reductions of 10–30% in optimized plants; lower material waste and fewer defective parts

Buildings & HVAC

Occupancy-based control, predictive climate control, fault detection

Heat/cool only occupied spaces; anticipate weather and usage; detect inefficient equipment early

Significant cuts in heating, cooling, and lighting energy (often 15–30% in smart buildings)

Transportation & Logistics

Route optimization, traffic management, demand prediction, autonomous systems

Reduce empty miles, congestion, and unnecessary trips; improve vehicle utilization and fuel efficiency

Lower fuel/energy use per ton-mile or passenger-mile; fewer vehicles needed for same service

Materials & Chemistry

Accelerated materials discovery and catalyst design (HPC simulation + ML)

Design better batteries, insulation, catalysts, and lightweight materials with fewer experiments

Higher-efficiency products (e.g., better batteries, lower-energy chemical processes) reduce lifetime resource intensity

Water systems

Leak detection, treatment optimization, demand forecasting

Identify and prioritize pipe leaks; optimize chemical dosing and energy in treatment plants

Reduced non-revenue water losses; lower energy and chemical use in water treatment

Supply chains & Circular economy

Demand forecasting, inventory optimization, computer-vision sorting for recycling

Cut overproduction and excess inventory; improve recovery of materials from waste streams

Less waste, lower storage/transport energy, higher material circularity


So a case can be made that current data center expansion concerns are akin to a developing moral panic.


Moral panics are:

  • marked by an increasing concern about a topic

  • characterized by a growing hostility toward the cause of the concern

  • marked by a consensus in the public debate about the nature of the threat

  • led by threats that are disproportionately larger than reality suggests.


Sociologist Stanley Cohen created the phrase. Moral panic is a widespread and exaggerated fear that an evil person, group, or entity threatens a community or society. 


Panic

Time Period

Underlying Cultural Anxiety

The "Folk Devil" or Target

The Salem Witch Trials

1692–1693

Religious anxiety, border warfare, and communal instability in colonial Massachusetts.

Marginalized women and community outsiders accused of witchcraft.

The Comic Book Panic

Late 1940s–1950s

Post-WWII anxiety over juvenile delinquency and the corruption of youth culture by mass media.

Comic book publishers (specifically horror and crime genres) and teenage readers.

Dungeons & Dragons and Satanic Panic

1980s

Fear of changing family structures, secularism, and the rise of youth fantasy subcultures.

Role-playing gamers, heavy metal music fans, and alleged secret satanic cults.

The "Super-Predator" Moral Panic

Mid-1990s

Fear of escalating urban crime rates and changing racial demographics in major cities.

Inner-city youth, particularly young Black males framed as remorseless criminals.


Such panics are relatively short-lived, but intense in the short term, in part because some have incentives to intensify the issue. Politicians; environmentalists; journalists and industry opponents are among them. 


The point is that, in such instances, the threats are exaggerated, often wildly. 


And if concern about data centers does develop into something like a moral panic, history suggests that concern will dissipate. 


The concerns are legitimate, but simply overblown.


Friday, August 21, 2026

Why "AI" is Not Like the Internet or Dot-Com Bubble

For some of us, analogies between the internet bubble around the turn of the century and a potential AI bubble often emphasize excesses of investment, but also questionable accounting practices and, in a few notable cases, outright fraud (Enron, Worldcom). 


If the basic AI market danger can be stated as overvaluation leading to overinvestment, creating financial stress that encourages aggressive accounting that then slips over to illegal actions, the main danger right now is still overinvestment. 


But accounting assumptions seem to raise some issues.


At least some observers of the high-performance computing industry and neocloud providers worry about possible financial excesses such as off-balance-sheet financing of graphics processing units; infrastructure overinvestment; circular financing and GPU depreciation assumptions. 


The legitimate concern is that demand will not ultimately support the supply, leading to a bubble collapse of firms and significant financial losses for investors. 


But accounting assumptions are among the contributing issues. The concern is that what is lawful might not be wise, at scale. 


But there might be some new information on GPU useful lives that allays some of the concern. Secondary market values of Nvidia H100 GPUs, an older generation, seem to be quite strong. 


NVIDIA H100 GPU prices in 2026 suggest that the cost of refurbished units is in the mid-80-percent range of new units. 


That is a  narrower discount than buyers expect from the refurbished category, where 30- to 50-percent discounts are typical. 


A mid-80-percent floor on a three-year-old accelerator suggests demand is strong enough that even second-hand units hold most of their value.


Older GPUs remain useful for operations other than frontier model operations. Even if the highest value for the latest generation of chips is to support frontier language model training, inference operations can still use older GPUs. Beyond that, many batch operations can be completed using processors that are five to six years old. 


So depreciation schedules embodying assumptions about six-year useful life are not an accounting trick. 


There are other users of such devices and chips as well. 


Still, there is some evidence that used GPU prices for the latest generations might depreciate faster than did older generations, as new generations are released faster.  


All that matters because depreciation assumptions bear directly on reported profits. 


If a GPU's true economic life is three years but that asset is depreciated over six years, the company understates depreciation expense and also overstates net income for years one to three.


It also then will take an accelerated depreciation later, which lowers reported income. 


Secondary market values for H100s seem to provide reassuring evidence that a six-year deprecation schedule is grounded in reality, and does not distort earnings. 


Still, some might worry about Enron-style excesses, but Enron’s accounting practices were not simply unwise, but unlawful. The same might be said of Worldcom.


Still, the main problems with the dot-com bubble relate to mistaken assumptions about demand, and subsequent oversupply. 


Question

Dot-com/Enron-era warning

AI equivalent

Is demand real?

Internet traffic was real, but forecasts became extreme

AI usage is clearly real—but is ultimate willingness to pay keeping pace with compute investment?

Does revenue come from outside the ecosystem?

Telecom companies sometimes effectively sold capacity to companies whose own economics depended on the same boom

Are AI companies buying from each other in ways that make industry revenue look larger than end-user demand?

Is infrastructure earning its cost of capital?

Fiber existed, but often couldn't generate adequate returns

Are GPUs/data centers/power assets generating sufficient cash flow over their useful lives?

Are accounting profits turning into cash?

Enron's mark-to-market profits could precede cash realization

Are AI-related profits accompanied by operating cash flow?

Are assets fairly valued?

Enron used models to value difficult-to-price assets

Are assumptions about GPU useful lives, residual values, utilization and AI infrastructure returns realistic?

Where is the debt?

Enron obscured liabilities through SPEs

Are AI infrastructure obligations sitting on balance sheets or in partnerships/project-finance structures?

Who ultimately bears the risk?

Financial structures redistributed risk

Who owns the downside if AI demand disappoints—AI developers, hyperscalers, chip companies, landlords, lenders or investors?

Is growth organic?

Acquisition and financial engineering could sustain reported growth

Are customers independently generating AI revenue, or is capital circulating among AI companies?

What happens if growth slows?

Small reductions in demand could make enormous infrastructure investments uneconomic

What happens if inference demand grows 30% instead of 100%?

Does valuation require perfection?

Dot-com valuations incorporated extraordinary future growth

What assumptions about revenue, margins and AI productivity are embedded in today's valuations?


To be sure, one resonant concern is the use of special purpose vehicles to move capital investment off balance sheets. 


To be fair, other capital-intensive industries, such as airlines, have used SPVs to finance aircraft. Power utilities use them for power plants. 


But it’s an area of concern. 


Circular transactions between value chain participants also are familiar issues. When the same $100 billion can show up as a chipmaker's revenue, a lab's funding, and a cloud's backlog, actual demand can be obscured.


On the other hand, Enron and Worldcom were guilty of outright fraud. Enron's core energy-trading business was dependent on accounting assumptions and actions. 


Nvidia, Microsoft, Amazon, Alphabet, and Meta have enormous real, profitable, non-AI-dependent businesses generating current cash flow.


The hyperscalers are unlikely to be in danger of an Enron-style collapse. But some neocloud providers without the existing cash flow and profits from other lines of business are at greater risk.

And that is why depreciation assumptions matter, especially for neocloud providers. 


But again, those assumptions ultimately matter only if demand does not develop as many expect. Yes, there are timing issues. 


Ideally, revenue scales in line with investment.


But it is ultimate demand that matters most, even if gross investment levels and payback timing also matter. 


Agentic AI Will Reshape Web Ad Economics

“For at least the last at least 30 years, the business model of the internet has been advertising,” says Matthew Prince, Cloudflare CEO . “I...