Skip to content
NYPTID Info

Encyclopedia article

Large language model

Neural network language models trained on large text corpora for language generation and related tasks

  • Updated
  • Version 1
  • 2,155 words
  • 9 min read
  • 19 sources
43 of 43 claims verified by two independent models
Share on Yap (opens in a new tab)

Report an issue

About this article.

What kind of issue?
Ask JRVS about this (opens in a new tab)

Truth Ledger

43 claims checked

Verified
43
Disputed
0
Removed
0

Checked by GPT-6.1 Sol and Grok 4.7, each without seeing the other’s answers.

Neutrality: 4 of 4 framing flags fixed.

A claim is stated as fact only when both checkers confirm it from the cited sources; a split verdict is published with attribution, and a claim neither can confirm is cut.

See every claim
Contents
  1. Overview
  2. Definition and scope
  3. History
  4. Training and architecture
  5. Capabilities and extensions
  6. Emergent abilities
  7. Evaluation
  8. Limitations and risks
  9. Debate over understanding
  10. Societal concerns
  11. Scripture
  12. Sources
  13. Truth Ledger

A large language model (LLM) is an artificial intelligence model, typically a neural network, trained on a vast amount of text for natural language processing tasks, especially language generation . LLMs can generate, summarize, translate and analyze text, and they underlie chatbots such as ChatGPT, Claude, Gemini, Grok and DeepSeek . Most are based on the transformer architecture, and they have drawn wide attention since the release of ChatGPT in November 2022 . Researchers disagree about whether LLMs understand language in any meaningful sense, and about how reliably benchmarks measure their abilities .

Definition and scope

#

Report an issue

In the section “Definition and scope”.

What kind of issue?

One survey describes LLMs as large-scale, pre-trained statistical language models based on neural networks . The same survey says the term mainly refers to transformer-based models with tens to hundreds of billions of parameters that are pre-trained on massive text data, such as PaLM, LLaMA and GPT-4 . It also notes that several non-transformer LLMs based on structured state space models have been proposed .

The term has no settled definition. The qualifier "large" has no definitive threshold for the number of parameters . A position paper asks how many parameters a network needs to qualify as an LLM, whether only transformer architectures qualify, and whether multimodal models such as text-to-image systems should be included . The same paper notes that the common technical definition of a language model is a model that assigns probabilities to upcoming words or word sequences, and that the term "large language model" is used differently by researchers, journalists and legislators .

In deployment, models are typically placed inside an external software harness or agent framework. This harness manages system instructions, tool access, memory and output formatting .

History

#

Report an issue

In the section “History”.

What kind of issue?

One survey dates language modeling to the 1950s, when Claude Shannon applied information theory to human language and measured how well simple n-gram models predict or compress text . The same survey groups later work into four waves: statistical language models, neural language models, pre-trained language models and LLMs . In the early 1990s, IBM's statistical models introduced word alignment techniques for machine translation . In 2001, a smoothed n-gram model trained on 300 million words achieved state-of-the-art perplexity on benchmark tests . Researchers began using neural networks as language models in 2000 .

After deep neural networks advanced image classification around 2012, similar architectures were adapted for language . Developments included word embeddings such as Word2Vec in 2013 and LSTM-based sequence-to-sequence models . In 2016, Google moved its translation service to neural machine translation .

Google researchers introduced the transformer architecture in the paper "Attention Is All You Need" . Sources disagree on the date. One places the paper at the 2017 NeurIPS conference . Microsoft's Azure documentation describes it as published in 2018 . BERT, an encoder-only model, was introduced in 2018 . Academic use of BERT began to decline in 2023, as decoder-only models such as GPT improved at solving tasks through prompting .

OpenAI introduced GPT-1 in 2018 . GPT-2 drew wide attention in 2019 after OpenAI said it had initially considered the model too powerful to release publicly because of possible malicious use . GPT-3 followed in 2020 . The consumer chatbot ChatGPT was released in late 2022 ; one survey dates its release to November 2022 . GPT-4 followed in 2023, and OpenAI did not disclose its architecture or parameter count .

In 2023, Meta's LLaMA developers reported that LLaMA-13B outperformed the 175-billion-parameter GPT-3 on most benchmarks . In 2024, OpenAI released the reasoning model o1, which generates long chains of thought before giving a final answer . In January 2025, DeepSeek released DeepSeek-R1, a 671-billion-parameter open-weight model . That source reports R1 performed comparably to o1 at a much lower price per token .

Training and architecture

#

Report an issue

In the section “Training and architecture”.

What kind of issue?

Text is first converted into numerical tokens, using algorithms such as byte-pair encoding and WordPiece . Training datasets are typically cleaned by removing duplicated data and material classified as low-quality or toxic under the developers' filtering criteria . Synthetic data may be used where naturally occurring data is insufficient in quantity or quality .

GPT models are first pretrained on large amounts of data to predict the next word, and are then fine-tuned . Fine-tuning shapes behavior through techniques such as reinforcement learning from human feedback (RLHF) and constitutional AI . In RLHF, a reward model is trained to predict which text humans prefer, and the LLM is then fine-tuned through reinforcement learning to better satisfy that reward model . In 2022, OpenAI demonstrated InstructGPT, a version of GPT-3 fine-tuned to follow instructions .

The transformer's attention mechanism lets a model process relationships between all elements of a sequence simultaneously . Autoregressive models such as GPT are trained to predict how a sequence continues . Masked models such as BERT are trained to predict missing parts of a sequence .

A mixture of experts (MoE) routes each input to specialized subnetworks, so only a fraction of the parameters is used per input, which can reduce inference costs . A 2025 review by Sebastian Raschka reports that open-weight LLMs largely converged on MoE layers that year . The same review says development in 2025 was dominated by reasoning models trained with reinforcement learning from verifiable rewards (RLVR) and group relative policy optimization (GRPO) .

Training the largest models requires substantial infrastructure . Reported training costs were $50,000 for GPT-2, a 1.5-billion-parameter model, in 2019; about $11 million for Megatron-Turing NLG 530B in 2021; and $8 million for PaLM, a 540-billion-parameter model, in 2022 . One billion parameters stored at 16-bit precision require 2 gigabytes . Post-training quantization lowers numerical precision to reduce storage needs while preserving most of a model's performance .

Chinchilla scaling, an empirical law, predicts pretraining loss from parameter count and training-set size . Under this formulation, training costs about 6 FLOPs per parameter per token, compared with 1 to 2 FLOPs per parameter per token for inference .

Capabilities and extensions

#

Report an issue

In the section “Capabilities and extensions”.

What kind of issue?

Few-shot prompting places worked examples in the input so that a model can be adapted to a task without fine-tuning . Chain-of-thought prompting has a model produce intermediate steps before its answer, and a 2022 paper found this improves correctness on relatively complex questions . Reasoning models are trained to generate step-by-step analysis before answering . On International Mathematics Olympiad qualifying exam problems, GPT-4o reportedly scored 13% accuracy and o1 scored 83% . Reasoning models typically need more computation per query than conventional LLMs .

Retrieval-augmented generation (RAG) retrieves documents relevant to a query and supplies them to the model as context . With tool use, a separate program watches the model's output for tool-calling syntax, executes the call, and feeds the result back to the model . An LLM is typically not an autonomous agent on its own, but it can function as one when memory, planning prompts and tool access are added .

Since 2023, many LLMs have been trained to be multimodal . A review in the Journal of Big Data cites GPT-4o and GPT-5 as examples of native vision-language understanding . The same review says Gemini 2.5 Pro offers context windows long enough to analyze hours of video . It also notes that LLaMA 3.2 added vision capabilities along with small 1B and 3B variants for local deployment .

LLMs handle programming languages much as they handle natural language, and services such as GitHub Copilot apply them to programming . Transformer-based models have also been applied to protein, DNA and RNA sequences . Meta's ESMFold, for example, predicts protein structure about an order of magnitude faster than AlphaFold2 . A structured review lists applications in healthcare, finance, education, agriculture, marketing, software engineering and scientific research .

Emergent abilities

#

Report an issue

In the section “Emergent abilities”.

What kind of issue?

One survey describes emergent abilities, absent in smaller models, that appear in LLMs . It names in-context learning, instruction following and multi-step reasoning . Other proposed examples include arithmetic, unscrambling words and decoding the International Phonetic Alphabet . A 2022 paper found that chain-of-thought prompting improved performance only for models with at least 62 billion parameters .

Schaeffer and colleagues dispute this framing. They argue that such abilities are acquired predictably, following a smooth scaling law, rather than appearing unpredictably .

Evaluation

#

Report an issue

In the section “Evaluation”.

What kind of issue?

The canonical measure of a language model's performance is perplexity on a text corpus . Benchmarks assess general knowledge, bias, commonsense reasoning, question answering and mathematics, and results are often sensitive to the prompting method . Adversarial datasets such as TruthfulQA, which contains 817 questions, target known failure modes .

Test-data contamination is a recognized problem. Larger models trained on larger corpora are increasingly likely to have seen portions of a test set during training . One position paper argues that most reported evaluation results should be treated cautiously because of contamination . It notes that GPT-4's reported top-10% score on a simulated bar exam was questioned on grounds of improper evaluation and possible contamination . The same paper reports that the LM Contamination Index then had 375 entries .

A reproducibility audit examined claims of "Potemkin understanding", in which a model correctly defines a concept but fails to apply it . Across five identical runs per model, the audit found reported Potemkin scores varied by as much as 31% . The audit also found that reasoning-focused models reduced measured incoherence by an order of magnitude . It concluded that claims of widespread conceptual incoherence in LLMs are "directionally supported but empirically fragile" .

Limitations and risks

#

Report an issue

In the section “Limitations and risks”.

What kind of issue?

Generative LLMs can confidently state claims that their training data do not appear to justify, a phenomenon termed hallucination . Efforts to reduce hallucination have used automated reasoning, RAG and fine-tuning . A Communications of the ACM article reports that these measures have made models far less likely to invent citations or statistics . The same article says models still lack an understanding of how the world works, and it uses the term "jagged intelligence" for their uneven performance . MIT's Phillip Isola is quoted there as saying that transformers do not natively have memory in the human sense . The article also states that GPT models consume enormous compute and energy resources .

LLMs can inherit and amplify biases in their training data, including gender, language and political biases . Because English dominates training corpora, models may favor English-language perspectives regardless of the language of the query . LLMs also tend toward sycophancy, producing responses they predict users want to hear rather than what is accurate .

Prompt injection lets users or third-party content cause a model to depart from its intended instructions or violate security controls, and LLMs have difficulty distinguishing user instructions from instructions embedded in web pages or files . Anthropic researchers showed that models could be trained as "sleeper agents" with hidden behaviors that safety training had difficulty removing . In 2025, the American Sunlight Project, a non-profit, reported evidence that the pro-Russia Pravda network was mass-publishing web content intended to bias LLM outputs .

Debate over understanding

#

Report an issue

In the section “Debate over understanding”.

What kind of issue?

In a 2022 survey, NLP researchers were evenly split on whether untuned LLMs could ever understand natural language in some nontrivial sense . A review in PNAS describes one group of researchers as arguing that these networks truly understand language and can reason in a general way, although not yet at human level . Another group cited in that review argues that LLMs likely capture important aspects of meaning through conceptual role .

Other researchers argue that models trained only to predict words learn the form of language but not its meaning, because they have no experience or mental models of the world . One scholar cited in the review describes LLMs as compressed repositories of human knowledge, closer to libraries than to intelligent agents . Some critics use the phrase "stochastic parrot" to argue that fluent language generation does not establish understanding .

The PNAS authors themselves propose an extended science of intelligence that would study distinct modes of understanding . A paper in PNAS Nexus identifies six misconceptions about LLMs, concerning next-token prediction, regression to the mean, training-data regurgitation, model memory, alignment and understanding .

In 2022, Google fired engineer Blake Lemoine, who had claimed its LaMDA model was conscious . Google described his claims as unfounded . There is no generally accepted method for determining whether an LLM is sentient .

Societal concerns

#

Report an issue

In the section “Societal concerns”.

What kind of issue?

LLMs sometimes reproduce long passages of training data verbatim, which raises copyright and memorization concerns . Web scraping to collect training data has caused denial-of-service problems for many websites .

Data centers used to train LLMs consume substantial electricity . A 2024 study by Luccioni, Jernite and Strubell measured an average of about 0.05 Wh per prompt for text generation and 2.91 Wh per prompt for image generation .

In early 2025, a Sentio University survey reported that 48.7% of 499 U.S. adults with ongoing mental health conditions who had used LLMs had turned to them for therapy or emotional support . This figure rests on a single survey, reported in one source . Researchers have also raised concerns that LLMs may lack effective crisis safety protocols .

Scripture

Passages quoted from the King James Version. The text is fetched, never written by a model.

My son, if thou wilt receive my words, and hide my commandments with thee; So that thou incline thine ear unto wisdom, and apply thine heart to understanding; Yea, if thou criest after knowledge, and liftest up thy voice for understanding; If thou seekest her as silver, and searchest for her as for hid treasures; Then shalt thou understand the fear of the LORD, and find the knowledge of God. For the LORD giveth wisdom: out of his mouth cometh knowledge and understanding.

Proverbs 2:1-6(King James Version)Real understanding is portrayed as something sought and received from the LORD, which bears on the debate over whether a text-trained model can truly understand or only process language.

He that answereth a matter before he heareth it, it is folly and shame unto him.

Proverbs 18:13(King James Version)Answering a matter before hearing it fully is called folly, a caution for people who rely on fluent machine-generated answers without checking them.

Prove all things; hold fast that which is good.

1 Thessalonians 5:21(King James Version)The command to prove all things and hold fast to what is good applies to testing the claims, outputs and benchmarks of AI systems.

And the whole earth was of one language, and of one speech. And it came to pass, as they journeyed from the east, that they found a plain in the land of Shinar; and they dwelt there. And they said one to another, Go to, let us make brick, and burn them throughly. And they had brick for stone, and slime had they for morter. And they said, Go to, let us build us a city and a tower, whose top may reach unto heaven; and let us make us a name, lest we be scattered abroad upon the face of the whole earth. And the LORD came down to see the city and the tower, which the children of men builded. And the LORD said, Behold, the people is one, and they have all one language; and this they begin to do: and now nothing will be restrained from them, which they have imagined to do.

Genesis 11:1-6(King James Version)The account of Babel shows human language and ingenuity used for ambitious projects, which invites reflection on the power and limits of language technology.

Sources

  1. 1.
    Large language model — Wikipedia (opens in a new tab)

    en.wikipedia.orgWikipedia (CC BY-SA 4.0)

  2. 2.
  3. 3.
  4. 4.
  5. 5.
  6. 6.
  7. 7.
  8. 8.
  9. 9.
  10. 10.
  11. 11.
  12. 12.
  13. 13.
  14. 14.
  15. 15.
  16. 16.
  17. 17.
  18. 18.
  19. 19.

Truth Ledger

Every checkable claim in the draft, checked by GPT-6.1 Sol and Grok 4.7. A claim is stated as fact only when both checkers confirm it from the cited sources; a split verdict is published with attribution, and a claim neither can confirm is cut.

Showing 43 claims.

  1. Verified

    A large language model is an AI model, typically a neural network, trained on vast amounts of text for natural language processing tasks, especially language generation.

    • Grok 4.7:Supported
    • GPT-6.1 Sol:Supported

    Source [1] explicitly gives this definition. / Source [1] defines an LLM in essentially these terms.

    Cites1

  2. Verified

    LLMs underlie chatbots including ChatGPT, Claude, Gemini, Grok and DeepSeek.

    • Grok 4.7:Supported
    • GPT-6.1 Sol:Supported

    Source [1] lists all five chatbots as based on LLMs. / Source [1] lists those chatbots as based on LLMs.

    Cites1

  3. Verified

    LLMs are typically based on the transformer architecture.

    • Grok 4.7:Supported
    • GPT-6.1 Sol:Supported

    Sources [1] and [14] identify transformers as the typical LLM architecture. / Sources [1] and [14] say LLMs are typically transformer-based.

    Cites114

  4. Verified

    There is no definitive parameter threshold for a language model to qualify as large.

    • Grok 4.7:Supported
    • GPT-6.1 Sol:Supported

    Source [1] explicitly states there is no definitive parameter threshold; [17] also highlights definitional uncertainty. / Source [1] says there is no definitive parameter threshold for 'large'.

    Cites117

  5. Verified

    Whether only transformer-type architectures and whether multimodal models qualify as LLMs is unsettled.

    • Grok 4.7:Supported
    • GPT-6.1 Sol:Supported

    Source [17] explicitly raises both architecture and multimodality as unresolved definitional questions. / Source [17] treats both qualification questions as unresolved definitional issues.

    Cites17

  6. Verified

    Language modeling dates to Shannon's application of information theory to language in the 1950s.

    • Grok 4.7:Supported
    • GPT-6.1 Sol:Supported

    Source [14] traces language modeling to Shannon’s information-theoretic work in the 1950s. / Source [14] dates language modeling to Shannon's 1950s work.

    Cites14

  7. Verified

    In 2001 a smoothed n-gram model trained on 300 million words achieved state-of-the-art perplexity.

    • Grok 4.7:Supported
    • GPT-6.1 Sol:Supported

    Source [1] gives the 2001 date, 300-million-word training corpus, and state-of-the-art perplexity result. / Source [1] states this 2001 n-gram result exactly.

    Cites1

  8. Verified

    Google moved its translation service to neural machine translation in 2016.

    • Grok 4.7:Supported
    • GPT-6.1 Sol:Supported

    Source [1] explicitly dates Google’s transition to neural machine translation to 2016. / Source [1] says Google moved translation to NMT in 2016.

    Cites1

  9. Verified

    The transformer was introduced in 'Attention Is All You Need' at the 2017 NeurIPS conference.

    • Grok 4.7:Supported
    • GPT-6.1 Sol:Supported

    Source [1] identifies the paper and its introduction at the 2017 NeurIPS conference. / Source [1] places the paper at NeurIPS 2017.

    Cites1

  10. Verified

    Microsoft's Azure documentation dates 'Attention Is All You Need' to 2018.

    • Grok 4.7:Supported
    • GPT-6.1 Sol:Supported

    Source [18] does date the paper to 2018, although the actual introduction was in 2017. / Source [18], Azure documentation, dates the paper to 2018.

    Cites18

  11. Verified

    BERT, an encoder-only model, was introduced in 2018.

    • Grok 4.7:Supported
    • GPT-6.1 Sol:Supported

    Source [1] states that BERT was introduced in 2018 and is encoder-only. / Source [1] says BERT was introduced in 2018 and is encoder-only.

    Cites1

  12. Verified

    OpenAI said it initially considered GPT-2 too powerful to release publicly for fear of malicious use.

    • Grok 4.7:Supported
    • GPT-6.1 Sol:Supported

    Source [1] explicitly attributes this initial release concern to OpenAI. / Source [1] attributes that claim to OpenAI about GPT-2.

    Cites1

  13. Verified

    ChatGPT was released in November 2022.

    • Grok 4.7:Supported
    • GPT-6.1 Sol:Supported

    Source [14] explicitly dates ChatGPT’s release to November 2022. / Source [14] dates ChatGPT's release to November 2022.

    Cites14

  14. Verified

    OpenAI did not disclose GPT-4's architecture or parameter count.

    • Grok 4.7:Supported
    • GPT-6.1 Sol:Supported

    Source [1] states that OpenAI did not reveal GPT-4’s high-level architecture or parameter count. / Source [1] says OpenAI did not reveal GPT-4's architecture or parameter count.

    Cites1

  15. Verified

    LLaMA's developers reported that LLaMA-13B outperformed GPT-3 (175B) on most benchmarks.

    • Grok 4.7:Supported
    • GPT-6.1 Sol:Supported

    Source [19] explicitly reports LLaMA-13B outperforming GPT-3 (175B) on most benchmarks. / Source [19] states LLaMA-13B outperforms GPT-3 (175B) on most benchmarks.

    Cites19

  16. Verified

    OpenAI released the reasoning model o1 in 2024.

    • Grok 4.7:Supported
    • GPT-6.1 Sol:Supported

    Source [1] dates o1’s release to 2024, specifically September. / Source [1] says OpenAI released o1 in 2024.

    Cites1

  17. Verified

    DeepSeek released DeepSeek-R1, a 671-billion-parameter open-weight model, in January 2025.

    • Grok 4.7:Supported
    • GPT-6.1 Sol:Supported

    Source [1] explicitly supplies the January 2025 release, 671-billion parameter count, and open-weight status. / Source [1] reports DeepSeek-R1's January 2025 release and 671B open weights.

    Cites1

  18. Verified

    In RLHF, a reward model is trained to predict human preferences and the LLM is fine-tuned with reinforcement learning against it.

    • Grok 4.7:Supported
    • GPT-6.1 Sol:Supported

    Source [1] describes reward-model preference prediction followed by reinforcement learning to satisfy that model. / Source [1] describes RLHF as preference reward modeling plus RL fine-tuning.

    Cites1

  19. Verified

    OpenAI demonstrated the instruction-tuned InstructGPT in 2022.

    • Grok 4.7:Supported
    • GPT-6.1 Sol:Supported

    Source [1] states OpenAI demonstrated instruction-tuned InstructGPT in 2022. / Source [1] says OpenAI demonstrated instruction-tuned InstructGPT in 2022.

    Cites1

  20. Verified

    Mixture-of-experts architectures use only a fraction of parameters per input and can reduce inference costs.

    • Grok 4.7:Supported
    • GPT-6.1 Sol:Supported

    Source [1] explicitly states that activating only a fraction of parameters can reduce inference costs. / Source [1] says MoE uses a fraction of parameters and can cut inference cost.

    Cites1

  21. Verified

    A 2025 review reports open-weight LLMs largely converged on mixture-of-experts layers that year.

    • Grok 4.7:Supported
    • GPT-6.1 Sol:Supported

    Source [2] reports that open-weight LLMs more or less converged on MoE layers in 2025. / Source [2] says open-weight LLMs more or less converged on MoE that year.

    Cites2

  22. Verified

    A 2025 review says LLM development that year was dominated by reasoning models trained with RLVR and GRPO.

    • Grok 4.7:Supported
    • GPT-6.1 Sol:Supported

    Source [2] explicitly characterizes 2025 development as dominated by reasoning models using RLVR and GRPO. / Source [2] says 2025 LLM development was dominated by reasoning models using RLVR and GRPO.

    Cites2

  23. Verified

    Training GPT-2 in 2019 cost $50,000 and training PaLM in 2022 cost $8 million.

    • Grok 4.7:Supported
    • GPT-6.1 Sol:Supported

    Source [1] gives both training-cost figures and their respective years. / Source [1] gives GPT-2's 2019 cost as $50,000 and PaLM's 2022 cost as $8 million.

    Cites1

  24. Verified

    Under Chinchilla scaling, training costs about 6 FLOPs per parameter per token versus 1 to 2 FLOPs for inference.

    • Grok 4.7:Supported
    • GPT-6.1 Sol:Supported

    Source [1] specifies six training FLOPs per parameter per token and one to two for inference. / Source [1] states 6 training FLOPs and 1–2 inference FLOPs per parameter per token.

    Cites1

  25. Verified

    On IMO qualifying exam problems GPT-4o reportedly scored 13% and o1 83%.

    • Grok 4.7:Supported
    • GPT-6.1 Sol:Supported

    Source [1] reports 13% accuracy for GPT-4o and 83% for o1 on IMO qualifying exam problems. / Source [1] reports 13% for GPT-4o and 83% for o1 on those problems.

    Cites1

  26. Verified

    LLaMA 3.2 introduced vision capabilities and 1B and 3B parameter variants for local deployment.

    • Grok 4.7:Supported
    • GPT-6.1 Sol:Supported

    Source [3] explicitly describes LLaMA 3.2’s vision capabilities and its 1B and 3B local-deployment variants. / Source [3] says LLaMA 3.2 added vision and 1B/3B local variants.

    Cites3

  27. Verified

    ESMFold predicts protein structure about an order of magnitude faster than AlphaFold2.

    • Grok 4.7:Supported
    • GPT-6.1 Sol:Supported

    Source [1] states that ESMFold runs an order of magnitude faster than AlphaFold2. / Source [1] says ESMFold runs an order of magnitude faster than AlphaFold2.

    Cites1

  28. Verified

    A 2022 paper found chain-of-thought prompting improved performance only in models with at least 62 billion parameters.

    • Grok 4.7:Supported
    • GPT-6.1 Sol:Supported

    Source [1] explicitly attributes this 62-billion-parameter threshold finding to a 2022 chain-of-thought paper. / Source [1] says a 2022 paper found CoT helped only models of at least 62B.

    Cites1

  29. Verified

    Schaeffer et al. argue emergent abilities are acquired predictably according to a smooth scaling law.

    • Grok 4.7:Supported
    • GPT-6.1 Sol:Supported

    Source [1] attributes predictable, smooth scaling of emergent abilities to Schaeffer and colleagues. / Source [1] attributes that smooth-scaling argument to Schaeffer et al.

    Cites1

  30. Verified

    TruthfulQA consists of 817 questions designed to elicit learned falsehoods.

    • Grok 4.7:Supported
    • GPT-6.1 Sol:Supported

    Source [1] describes TruthfulQA as 817 questions targeting falsehoods encountered during training. / Source [1] describes TruthfulQA as 817 questions mimicking trained falsehoods.

    Cites1

  31. Verified

    GPT-4's reported top-10% simulated bar exam score was questioned on grounds of improper evaluation and possible data contamination.

    • Grok 4.7:Supported
    • GPT-6.1 Sol:Supported

    Source [17] explicitly describes these objections to GPT-4’s reported simulated bar exam ranking. / Source [17] says the top-10% bar-exam claim was questioned for evaluation and contamination.

    Cites17

  32. Verified

    A reproducibility audit found Potemkin scores varied by as much as 31% across five identical runs per model.

    • Grok 4.7:Supported
    • GPT-6.1 Sol:Supported

    Source [9] reports variation of up to 31% across five identical runs per model. / Source [9] reports Potemkin scores varying by up to 31% across five runs.

    Cites9

  33. Verified

    The audit found reasoning-focused models reduced measured incoherence by an order of magnitude.

    • Grok 4.7:Supported
    • GPT-6.1 Sol:Supported

    Source [9] explicitly reports an order-of-magnitude reduction in incoherence for reasoning-focused models. / Source [9] says reasoning-focused models reduced incoherence by an order of magnitude.

    Cites9

  34. Verified

    A Communications of the ACM article reports that LLMs are far less likely than before to invent citations or statistics but still lack an understanding of how the world works.

    • Grok 4.7:Supported
    • GPT-6.1 Sol:Supported

    Source [10] makes both statements; the claim accurately attributes them to the article rather than asserting scientific consensus. / Source [10] says invented citations/statistics are far less likely, but world understanding remains lacking.

    Cites10

  35. Verified

    NLP researchers were evenly split in a 2022 survey on whether untuned LLMs could ever understand language in a nontrivial sense.

    • Grok 4.7:Supported
    • GPT-6.1 Sol:Supported

    Source [1] reports an even split in the 2022 survey on untuned LLMs’ potential language understanding. / Source [1] reports an even 2022 split on nontrivial understanding by untuned LLMs.

    Cites1

  36. Verified

    Some researchers argue LLMs learn the form of language but not its meaning because they lack experience or world models.

    • Grok 4.7:Supported
    • GPT-6.1 Sol:Supported

    Source [4] explicitly presents this argument as one side of the understanding debate. / Source [4] says critics argue text-only training yields form without meaning or world models.

    Cites4

  37. Verified

    A PNAS Nexus paper identifies six misconceptions about LLMs, including next-token prediction, alignment and understanding.

    • Grok 4.7:Supported
    • GPT-6.1 Sol:Supported

    Source [5] lists six misconceptions, including next-token prediction, alignment, and understanding. / Source [5] lists six misconceptions, including next-token prediction, alignment, and understanding.

    Cites5

  38. Verified

    Google fired engineer Blake Lemoine in 2022 after he claimed LaMDA was conscious, and described his claims as unfounded.

    • Grok 4.7:Supported
    • GPT-6.1 Sol:Supported

    Source [1] explicitly gives the firing year, Lemoine’s consciousness claim, and Google’s rejection of it as unfounded. / Source [1] says Google fired Lemoine in 2022 and called his LaMDA claims unfounded.

    Cites1

  39. Verified

    Anthropic researchers showed models could be trained as sleeper agents whose hidden behaviors were difficult to remove via safety training.

    • Grok 4.7:Supported
    • GPT-6.1 Sol:Supported

    Source [1] describes Anthropic’s sleeper-agent experiments and the difficulty of removing their hidden behaviors through safety training. / Source [1] says Anthropic created sleeper agents whose behaviors were hard to remove by safety training.

    Cites1

  40. Verified

    The American Sunlight Project reported in 2025 that the Pravda network was mass-publishing content intended to bias LLM outputs.

    • Grok 4.7:Supported
    • GPT-6.1 Sol:Supported

    Source [1] describes the 2025 report, Pravda network, mass publication, and intended biasing of LLM outputs. / Source [1] says the 2025 study found Pravda mass-publishing to bias LLM outputs.

    Cites1

  41. Verified

    Luccioni, Jernite and Strubell measured about 0.05 Wh per prompt for text generation and 2.91 Wh per prompt for image generation.

    • Grok 4.7:Supported
    • GPT-6.1 Sol:Supported

    Source [1] attributes these average per-prompt energy figures to Luccioni, Jernite, and Strubell’s 2024 study. / Source [1] cites their study: about 0.05 Wh for text and 2.91 Wh for images.

    Cites1

  42. Verified

    A 2025 Sentio University survey found 48.7% of 499 U.S. adults with ongoing mental health conditions who used LLMs turned to them for therapy or emotional support.

    • Grok 4.7:Supported
    • GPT-6.1 Sol:Supported

    Source [1] gives the 2025 survey, 499-person population, 48.7% figure, and therapy or emotional-support use. / Source [1] reports that 48.7% figure from the early-2025 Sentio survey of 499 adults.

    Cites1

  43. Verified

    Web scraping for LLM training data has caused denial-of-service problems for many websites.

    • Grok 4.7:Supported
    • GPT-6.1 Sol:Supported

    Source [1] explicitly links LLM training-data scraping traffic to denial-of-service problems affecting many websites. / Source [1] says LLM training scrapers caused denial-of-service issues for many websites.

    Cites1

Text is available under the Creative Commons Attribution-ShareAlike 4.0 licence. Written by Claude Opus 5.5 from the sources listed and checked claim by claim by GPT-6.1 Sol and Grok 4.7.