Encyclopedia article
Large language model
Neural network language models trained on large text corpora for language generation and related tasks
- Updated
- Version 1
- 2,155 words
- 9 min read
- 19 sources
Truth Ledger
43 claims checked
- Verified
- 43
- Disputed
- 0
- Removed
- 0
Checked by GPT-6.1 Sol and Grok 4.7, each without seeing the other’s answers.
Neutrality: 4 of 4 framing flags fixed.
A claim is stated as fact only when both checkers confirm it from the cited sources; a split verdict is published with attribution, and a claim neither can confirm is cut.
See every claimA large language model (LLM) is an artificial intelligence model, typically a neural network, trained on a vast amount of text for natural language processing tasks, especially language generation 1. LLMs can generate, summarize, translate and analyze text, and they underlie chatbots such as ChatGPT, Claude, Gemini, Grok and DeepSeek 1. Most are based on the transformer architecture, and they have drawn wide attention since the release of ChatGPT in November 2022 114. Researchers disagree about whether LLMs understand language in any meaningful sense, and about how reliably benchmarks measure their abilities 14917.
Definition and scope
#One survey describes LLMs as large-scale, pre-trained statistical language models based on neural networks 14. The same survey says the term mainly refers to transformer-based models with tens to hundreds of billions of parameters that are pre-trained on massive text data, such as PaLM, LLaMA and GPT-4 14. It also notes that several non-transformer LLMs based on structured state space models have been proposed 14.
The term has no settled definition. The qualifier "large" has no definitive threshold for the number of parameters 1. A position paper asks how many parameters a network needs to qualify as an LLM, whether only transformer architectures qualify, and whether multimodal models such as text-to-image systems should be included 17. The same paper notes that the common technical definition of a language model is a model that assigns probabilities to upcoming words or word sequences, and that the term "large language model" is used differently by researchers, journalists and legislators 17.
In deployment, models are typically placed inside an external software harness or agent framework. This harness manages system instructions, tool access, memory and output formatting 1.
History
#One survey dates language modeling to the 1950s, when Claude Shannon applied information theory to human language and measured how well simple n-gram models predict or compress text 14. The same survey groups later work into four waves: statistical language models, neural language models, pre-trained language models and LLMs 14. In the early 1990s, IBM's statistical models introduced word alignment techniques for machine translation 1. In 2001, a smoothed n-gram model trained on 300 million words achieved state-of-the-art perplexity on benchmark tests 1. Researchers began using neural networks as language models in 2000 1.
After deep neural networks advanced image classification around 2012, similar architectures were adapted for language 1. Developments included word embeddings such as Word2Vec in 2013 and LSTM-based sequence-to-sequence models 1. In 2016, Google moved its translation service to neural machine translation 1.
Google researchers introduced the transformer architecture in the paper "Attention Is All You Need" 118. Sources disagree on the date. One places the paper at the 2017 NeurIPS conference 1. Microsoft's Azure documentation describes it as published in 2018 18. BERT, an encoder-only model, was introduced in 2018 1. Academic use of BERT began to decline in 2023, as decoder-only models such as GPT improved at solving tasks through prompting 1.
OpenAI introduced GPT-1 in 2018 1. GPT-2 drew wide attention in 2019 after OpenAI said it had initially considered the model too powerful to release publicly because of possible malicious use 1. GPT-3 followed in 2020 1. The consumer chatbot ChatGPT was released in late 2022 1; one survey dates its release to November 2022 14. GPT-4 followed in 2023, and OpenAI did not disclose its architecture or parameter count 1.
In 2023, Meta's LLaMA developers reported that LLaMA-13B outperformed the 175-billion-parameter GPT-3 on most benchmarks 19. In 2024, OpenAI released the reasoning model o1, which generates long chains of thought before giving a final answer 1. In January 2025, DeepSeek released DeepSeek-R1, a 671-billion-parameter open-weight model 1. That source reports R1 performed comparably to o1 at a much lower price per token 1.
Training and architecture
#Text is first converted into numerical tokens, using algorithms such as byte-pair encoding and WordPiece 1. Training datasets are typically cleaned by removing duplicated data and material classified as low-quality or toxic under the developers' filtering criteria 1. Synthetic data may be used where naturally occurring data is insufficient in quantity or quality 1.
GPT models are first pretrained on large amounts of data to predict the next word, and are then fine-tuned 1. Fine-tuning shapes behavior through techniques such as reinforcement learning from human feedback (RLHF) and constitutional AI 1. In RLHF, a reward model is trained to predict which text humans prefer, and the LLM is then fine-tuned through reinforcement learning to better satisfy that reward model 1. In 2022, OpenAI demonstrated InstructGPT, a version of GPT-3 fine-tuned to follow instructions 1.
The transformer's attention mechanism lets a model process relationships between all elements of a sequence simultaneously 1. Autoregressive models such as GPT are trained to predict how a sequence continues 1. Masked models such as BERT are trained to predict missing parts of a sequence 1.
A mixture of experts (MoE) routes each input to specialized subnetworks, so only a fraction of the parameters is used per input, which can reduce inference costs 1. A 2025 review by Sebastian Raschka reports that open-weight LLMs largely converged on MoE layers that year 2. The same review says development in 2025 was dominated by reasoning models trained with reinforcement learning from verifiable rewards (RLVR) and group relative policy optimization (GRPO) 2.
Training the largest models requires substantial infrastructure 1. Reported training costs were $50,000 for GPT-2, a 1.5-billion-parameter model, in 2019; about $11 million for Megatron-Turing NLG 530B in 2021; and $8 million for PaLM, a 540-billion-parameter model, in 2022 1. One billion parameters stored at 16-bit precision require 2 gigabytes 1. Post-training quantization lowers numerical precision to reduce storage needs while preserving most of a model's performance 1.
Chinchilla scaling, an empirical law, predicts pretraining loss from parameter count and training-set size 1. Under this formulation, training costs about 6 FLOPs per parameter per token, compared with 1 to 2 FLOPs per parameter per token for inference 1.
Capabilities and extensions
#Few-shot prompting places worked examples in the input so that a model can be adapted to a task without fine-tuning 1. Chain-of-thought prompting has a model produce intermediate steps before its answer, and a 2022 paper found this improves correctness on relatively complex questions 1. Reasoning models are trained to generate step-by-step analysis before answering 1. On International Mathematics Olympiad qualifying exam problems, GPT-4o reportedly scored 13% accuracy and o1 scored 83% 1. Reasoning models typically need more computation per query than conventional LLMs 1.
Retrieval-augmented generation (RAG) retrieves documents relevant to a query and supplies them to the model as context 1. With tool use, a separate program watches the model's output for tool-calling syntax, executes the call, and feeds the result back to the model 1. An LLM is typically not an autonomous agent on its own, but it can function as one when memory, planning prompts and tool access are added 1.
Since 2023, many LLMs have been trained to be multimodal 1. A review in the Journal of Big Data cites GPT-4o and GPT-5 as examples of native vision-language understanding 3. The same review says Gemini 2.5 Pro offers context windows long enough to analyze hours of video 3. It also notes that LLaMA 3.2 added vision capabilities along with small 1B and 3B variants for local deployment 3.
LLMs handle programming languages much as they handle natural language, and services such as GitHub Copilot apply them to programming 1. Transformer-based models have also been applied to protein, DNA and RNA sequences 1. Meta's ESMFold, for example, predicts protein structure about an order of magnitude faster than AlphaFold2 1. A structured review lists applications in healthcare, finance, education, agriculture, marketing, software engineering and scientific research 6.
Emergent abilities
#One survey describes emergent abilities, absent in smaller models, that appear in LLMs 14. It names in-context learning, instruction following and multi-step reasoning 14. Other proposed examples include arithmetic, unscrambling words and decoding the International Phonetic Alphabet 1. A 2022 paper found that chain-of-thought prompting improved performance only for models with at least 62 billion parameters 1.
Schaeffer and colleagues dispute this framing. They argue that such abilities are acquired predictably, following a smooth scaling law, rather than appearing unpredictably 1.
Evaluation
#The canonical measure of a language model's performance is perplexity on a text corpus 1. Benchmarks assess general knowledge, bias, commonsense reasoning, question answering and mathematics, and results are often sensitive to the prompting method 1. Adversarial datasets such as TruthfulQA, which contains 817 questions, target known failure modes 1.
Test-data contamination is a recognized problem. Larger models trained on larger corpora are increasingly likely to have seen portions of a test set during training 1. One position paper argues that most reported evaluation results should be treated cautiously because of contamination 17. It notes that GPT-4's reported top-10% score on a simulated bar exam was questioned on grounds of improper evaluation and possible contamination 17. The same paper reports that the LM Contamination Index then had 375 entries 17.
A reproducibility audit examined claims of "Potemkin understanding", in which a model correctly defines a concept but fails to apply it 9. Across five identical runs per model, the audit found reported Potemkin scores varied by as much as 31% 9. The audit also found that reasoning-focused models reduced measured incoherence by an order of magnitude 9. It concluded that claims of widespread conceptual incoherence in LLMs are "directionally supported but empirically fragile" 9.
Limitations and risks
#Generative LLMs can confidently state claims that their training data do not appear to justify, a phenomenon termed hallucination 1. Efforts to reduce hallucination have used automated reasoning, RAG and fine-tuning 1. A Communications of the ACM article reports that these measures have made models far less likely to invent citations or statistics 10. The same article says models still lack an understanding of how the world works, and it uses the term "jagged intelligence" for their uneven performance 10. MIT's Phillip Isola is quoted there as saying that transformers do not natively have memory in the human sense 10. The article also states that GPT models consume enormous compute and energy resources 10.
LLMs can inherit and amplify biases in their training data, including gender, language and political biases 1. Because English dominates training corpora, models may favor English-language perspectives regardless of the language of the query 1. LLMs also tend toward sycophancy, producing responses they predict users want to hear rather than what is accurate 1.
Prompt injection lets users or third-party content cause a model to depart from its intended instructions or violate security controls, and LLMs have difficulty distinguishing user instructions from instructions embedded in web pages or files 1. Anthropic researchers showed that models could be trained as "sleeper agents" with hidden behaviors that safety training had difficulty removing 1. In 2025, the American Sunlight Project, a non-profit, reported evidence that the pro-Russia Pravda network was mass-publishing web content intended to bias LLM outputs 1.
Debate over understanding
#In a 2022 survey, NLP researchers were evenly split on whether untuned LLMs could ever understand natural language in some nontrivial sense 1. A review in PNAS describes one group of researchers as arguing that these networks truly understand language and can reason in a general way, although not yet at human level 4. Another group cited in that review argues that LLMs likely capture important aspects of meaning through conceptual role 4.
Other researchers argue that models trained only to predict words learn the form of language but not its meaning, because they have no experience or mental models of the world 4. One scholar cited in the review describes LLMs as compressed repositories of human knowledge, closer to libraries than to intelligent agents 4. Some critics use the phrase "stochastic parrot" to argue that fluent language generation does not establish understanding 1.
The PNAS authors themselves propose an extended science of intelligence that would study distinct modes of understanding 4. A paper in PNAS Nexus identifies six misconceptions about LLMs, concerning next-token prediction, regression to the mean, training-data regurgitation, model memory, alignment and understanding 5.
In 2022, Google fired engineer Blake Lemoine, who had claimed its LaMDA model was conscious 1. Google described his claims as unfounded 1. There is no generally accepted method for determining whether an LLM is sentient 1.
Societal concerns
#LLMs sometimes reproduce long passages of training data verbatim, which raises copyright and memorization concerns 1. Web scraping to collect training data has caused denial-of-service problems for many websites 1.
Data centers used to train LLMs consume substantial electricity 1. A 2024 study by Luccioni, Jernite and Strubell measured an average of about 0.05 Wh per prompt for text generation and 2.91 Wh per prompt for image generation 1.
In early 2025, a Sentio University survey reported that 48.7% of 499 U.S. adults with ongoing mental health conditions who had used LLMs had turned to them for therapy or emotional support 1. This figure rests on a single survey, reported in one source 1. Researchers have also raised concerns that LLMs may lack effective crisis safety protocols 1.
Scripture
Passages quoted from the King James Version. The text is fetched, never written by a model.
My son, if thou wilt receive my words, and hide my commandments with thee; So that thou incline thine ear unto wisdom, and apply thine heart to understanding; Yea, if thou criest after knowledge, and liftest up thy voice for understanding; If thou seekest her as silver, and searchest for her as for hid treasures; Then shalt thou understand the fear of the LORD, and find the knowledge of God. For the LORD giveth wisdom: out of his mouth cometh knowledge and understanding.
He that answereth a matter before he heareth it, it is folly and shame unto him.
Prove all things; hold fast that which is good.
And the whole earth was of one language, and of one speech. And it came to pass, as they journeyed from the east, that they found a plain in the land of Shinar; and they dwelt there. And they said one to another, Go to, let us make brick, and burn them throughly. And they had brick for stone, and slime had they for morter. And they said, Go to, let us build us a city and a tower, whose top may reach unto heaven; and let us make us a name, lest we be scattered abroad upon the face of the whole earth. And the LORD came down to see the city and the tower, which the children of men builded. And the LORD said, Behold, the people is one, and they have all one language; and this they begin to do: and now nothing will be restrained from them, which they have imagined to do.
Sources
- 1.Large language model — Wikipedia (opens in a new tab)
en.wikipedia.orgWikipedia (CC BY-SA 4.0)
- 2.The State Of LLMs 2025: Progress, Problems, and Predictions (opens in a new tab)
magazine.sebastianraschka.com
- 3.
- 4.The debate over understanding in AI's large language models (opens in a new tab)
pmc.ncbi.nlm.nih.gov
- 5.
- 6.
- 7.
- 8.
- 9.
- 10.
- 11.
- 12.
- 13.History of Large Language Models (opens in a new tab)
link.springer.com
- 14.
- 15.
- 16.
- 17.arxiv.org (opens in a new tab)
arxiv.org
- 18.What are large language models (LLMs)? (opens in a new tab)
azure.microsoft.com
- 19.
Truth Ledger
Every checkable claim in the draft, checked by GPT-6.1 Sol and Grok 4.7. A claim is stated as fact only when both checkers confirm it from the cited sources; a split verdict is published with attribution, and a claim neither can confirm is cut.
Showing 43 claims.
- Verified
A large language model is an AI model, typically a neural network, trained on vast amounts of text for natural language processing tasks, especially language generation.
- Grok 4.7:Supported
- GPT-6.1 Sol:Supported
Source [1] explicitly gives this definition. / Source [1] defines an LLM in essentially these terms.
Cites1
- Verified
LLMs underlie chatbots including ChatGPT, Claude, Gemini, Grok and DeepSeek.
- Grok 4.7:Supported
- GPT-6.1 Sol:Supported
Source [1] lists all five chatbots as based on LLMs. / Source [1] lists those chatbots as based on LLMs.
Cites1
- Verified
There is no definitive parameter threshold for a language model to qualify as large.
- Grok 4.7:Supported
- GPT-6.1 Sol:Supported
Source [1] explicitly states there is no definitive parameter threshold; [17] also highlights definitional uncertainty. / Source [1] says there is no definitive parameter threshold for 'large'.
- Verified
Whether only transformer-type architectures and whether multimodal models qualify as LLMs is unsettled.
- Grok 4.7:Supported
- GPT-6.1 Sol:Supported
Source [17] explicitly raises both architecture and multimodality as unresolved definitional questions. / Source [17] treats both qualification questions as unresolved definitional issues.
Cites17
- Verified
Language modeling dates to Shannon's application of information theory to language in the 1950s.
- Grok 4.7:Supported
- GPT-6.1 Sol:Supported
Source [14] traces language modeling to Shannon’s information-theoretic work in the 1950s. / Source [14] dates language modeling to Shannon's 1950s work.
Cites14
- Verified
In 2001 a smoothed n-gram model trained on 300 million words achieved state-of-the-art perplexity.
- Grok 4.7:Supported
- GPT-6.1 Sol:Supported
Source [1] gives the 2001 date, 300-million-word training corpus, and state-of-the-art perplexity result. / Source [1] states this 2001 n-gram result exactly.
Cites1
- Verified
Google moved its translation service to neural machine translation in 2016.
- Grok 4.7:Supported
- GPT-6.1 Sol:Supported
Source [1] explicitly dates Google’s transition to neural machine translation to 2016. / Source [1] says Google moved translation to NMT in 2016.
Cites1
- Verified
The transformer was introduced in 'Attention Is All You Need' at the 2017 NeurIPS conference.
- Grok 4.7:Supported
- GPT-6.1 Sol:Supported
Source [1] identifies the paper and its introduction at the 2017 NeurIPS conference. / Source [1] places the paper at NeurIPS 2017.
Cites1
- Verified
Microsoft's Azure documentation dates 'Attention Is All You Need' to 2018.
- Grok 4.7:Supported
- GPT-6.1 Sol:Supported
Source [18] does date the paper to 2018, although the actual introduction was in 2017. / Source [18], Azure documentation, dates the paper to 2018.
Cites18
- Verified
BERT, an encoder-only model, was introduced in 2018.
- Grok 4.7:Supported
- GPT-6.1 Sol:Supported
Source [1] states that BERT was introduced in 2018 and is encoder-only. / Source [1] says BERT was introduced in 2018 and is encoder-only.
Cites1
- Verified
OpenAI said it initially considered GPT-2 too powerful to release publicly for fear of malicious use.
- Grok 4.7:Supported
- GPT-6.1 Sol:Supported
Source [1] explicitly attributes this initial release concern to OpenAI. / Source [1] attributes that claim to OpenAI about GPT-2.
Cites1
- Verified
ChatGPT was released in November 2022.
- Grok 4.7:Supported
- GPT-6.1 Sol:Supported
Source [14] explicitly dates ChatGPT’s release to November 2022. / Source [14] dates ChatGPT's release to November 2022.
Cites14
- Verified
OpenAI did not disclose GPT-4's architecture or parameter count.
- Grok 4.7:Supported
- GPT-6.1 Sol:Supported
Source [1] states that OpenAI did not reveal GPT-4’s high-level architecture or parameter count. / Source [1] says OpenAI did not reveal GPT-4's architecture or parameter count.
Cites1
- Verified
LLaMA's developers reported that LLaMA-13B outperformed GPT-3 (175B) on most benchmarks.
- Grok 4.7:Supported
- GPT-6.1 Sol:Supported
Source [19] explicitly reports LLaMA-13B outperforming GPT-3 (175B) on most benchmarks. / Source [19] states LLaMA-13B outperforms GPT-3 (175B) on most benchmarks.
Cites19
- Verified
OpenAI released the reasoning model o1 in 2024.
- Grok 4.7:Supported
- GPT-6.1 Sol:Supported
Source [1] dates o1’s release to 2024, specifically September. / Source [1] says OpenAI released o1 in 2024.
Cites1
- Verified
DeepSeek released DeepSeek-R1, a 671-billion-parameter open-weight model, in January 2025.
- Grok 4.7:Supported
- GPT-6.1 Sol:Supported
Source [1] explicitly supplies the January 2025 release, 671-billion parameter count, and open-weight status. / Source [1] reports DeepSeek-R1's January 2025 release and 671B open weights.
Cites1
- Verified
In RLHF, a reward model is trained to predict human preferences and the LLM is fine-tuned with reinforcement learning against it.
- Grok 4.7:Supported
- GPT-6.1 Sol:Supported
Source [1] describes reward-model preference prediction followed by reinforcement learning to satisfy that model. / Source [1] describes RLHF as preference reward modeling plus RL fine-tuning.
Cites1
- Verified
OpenAI demonstrated the instruction-tuned InstructGPT in 2022.
- Grok 4.7:Supported
- GPT-6.1 Sol:Supported
Source [1] states OpenAI demonstrated instruction-tuned InstructGPT in 2022. / Source [1] says OpenAI demonstrated instruction-tuned InstructGPT in 2022.
Cites1
- Verified
Mixture-of-experts architectures use only a fraction of parameters per input and can reduce inference costs.
- Grok 4.7:Supported
- GPT-6.1 Sol:Supported
Source [1] explicitly states that activating only a fraction of parameters can reduce inference costs. / Source [1] says MoE uses a fraction of parameters and can cut inference cost.
Cites1
- Verified
A 2025 review reports open-weight LLMs largely converged on mixture-of-experts layers that year.
- Grok 4.7:Supported
- GPT-6.1 Sol:Supported
Source [2] reports that open-weight LLMs more or less converged on MoE layers in 2025. / Source [2] says open-weight LLMs more or less converged on MoE that year.
Cites2
- Verified
A 2025 review says LLM development that year was dominated by reasoning models trained with RLVR and GRPO.
- Grok 4.7:Supported
- GPT-6.1 Sol:Supported
Source [2] explicitly characterizes 2025 development as dominated by reasoning models using RLVR and GRPO. / Source [2] says 2025 LLM development was dominated by reasoning models using RLVR and GRPO.
Cites2
- Verified
Training GPT-2 in 2019 cost $50,000 and training PaLM in 2022 cost $8 million.
- Grok 4.7:Supported
- GPT-6.1 Sol:Supported
Source [1] gives both training-cost figures and their respective years. / Source [1] gives GPT-2's 2019 cost as $50,000 and PaLM's 2022 cost as $8 million.
Cites1
- Verified
Under Chinchilla scaling, training costs about 6 FLOPs per parameter per token versus 1 to 2 FLOPs for inference.
- Grok 4.7:Supported
- GPT-6.1 Sol:Supported
Source [1] specifies six training FLOPs per parameter per token and one to two for inference. / Source [1] states 6 training FLOPs and 1–2 inference FLOPs per parameter per token.
Cites1
- Verified
On IMO qualifying exam problems GPT-4o reportedly scored 13% and o1 83%.
- Grok 4.7:Supported
- GPT-6.1 Sol:Supported
Source [1] reports 13% accuracy for GPT-4o and 83% for o1 on IMO qualifying exam problems. / Source [1] reports 13% for GPT-4o and 83% for o1 on those problems.
Cites1
- Verified
LLaMA 3.2 introduced vision capabilities and 1B and 3B parameter variants for local deployment.
- Grok 4.7:Supported
- GPT-6.1 Sol:Supported
Source [3] explicitly describes LLaMA 3.2’s vision capabilities and its 1B and 3B local-deployment variants. / Source [3] says LLaMA 3.2 added vision and 1B/3B local variants.
Cites3
- Verified
ESMFold predicts protein structure about an order of magnitude faster than AlphaFold2.
- Grok 4.7:Supported
- GPT-6.1 Sol:Supported
Source [1] states that ESMFold runs an order of magnitude faster than AlphaFold2. / Source [1] says ESMFold runs an order of magnitude faster than AlphaFold2.
Cites1
- Verified
A 2022 paper found chain-of-thought prompting improved performance only in models with at least 62 billion parameters.
- Grok 4.7:Supported
- GPT-6.1 Sol:Supported
Source [1] explicitly attributes this 62-billion-parameter threshold finding to a 2022 chain-of-thought paper. / Source [1] says a 2022 paper found CoT helped only models of at least 62B.
Cites1
- Verified
Schaeffer et al. argue emergent abilities are acquired predictably according to a smooth scaling law.
- Grok 4.7:Supported
- GPT-6.1 Sol:Supported
Source [1] attributes predictable, smooth scaling of emergent abilities to Schaeffer and colleagues. / Source [1] attributes that smooth-scaling argument to Schaeffer et al.
Cites1
- Verified
TruthfulQA consists of 817 questions designed to elicit learned falsehoods.
- Grok 4.7:Supported
- GPT-6.1 Sol:Supported
Source [1] describes TruthfulQA as 817 questions targeting falsehoods encountered during training. / Source [1] describes TruthfulQA as 817 questions mimicking trained falsehoods.
Cites1
- Verified
GPT-4's reported top-10% simulated bar exam score was questioned on grounds of improper evaluation and possible data contamination.
- Grok 4.7:Supported
- GPT-6.1 Sol:Supported
Source [17] explicitly describes these objections to GPT-4’s reported simulated bar exam ranking. / Source [17] says the top-10% bar-exam claim was questioned for evaluation and contamination.
Cites17
- Verified
A reproducibility audit found Potemkin scores varied by as much as 31% across five identical runs per model.
- Grok 4.7:Supported
- GPT-6.1 Sol:Supported
Source [9] reports variation of up to 31% across five identical runs per model. / Source [9] reports Potemkin scores varying by up to 31% across five runs.
Cites9
- Verified
The audit found reasoning-focused models reduced measured incoherence by an order of magnitude.
- Grok 4.7:Supported
- GPT-6.1 Sol:Supported
Source [9] explicitly reports an order-of-magnitude reduction in incoherence for reasoning-focused models. / Source [9] says reasoning-focused models reduced incoherence by an order of magnitude.
Cites9
- Verified
A Communications of the ACM article reports that LLMs are far less likely than before to invent citations or statistics but still lack an understanding of how the world works.
- Grok 4.7:Supported
- GPT-6.1 Sol:Supported
Source [10] makes both statements; the claim accurately attributes them to the article rather than asserting scientific consensus. / Source [10] says invented citations/statistics are far less likely, but world understanding remains lacking.
Cites10
- Verified
NLP researchers were evenly split in a 2022 survey on whether untuned LLMs could ever understand language in a nontrivial sense.
- Grok 4.7:Supported
- GPT-6.1 Sol:Supported
Source [1] reports an even split in the 2022 survey on untuned LLMs’ potential language understanding. / Source [1] reports an even 2022 split on nontrivial understanding by untuned LLMs.
Cites1
- Verified
Some researchers argue LLMs learn the form of language but not its meaning because they lack experience or world models.
- Grok 4.7:Supported
- GPT-6.1 Sol:Supported
Source [4] explicitly presents this argument as one side of the understanding debate. / Source [4] says critics argue text-only training yields form without meaning or world models.
Cites4
- Verified
A PNAS Nexus paper identifies six misconceptions about LLMs, including next-token prediction, alignment and understanding.
- Grok 4.7:Supported
- GPT-6.1 Sol:Supported
Source [5] lists six misconceptions, including next-token prediction, alignment, and understanding. / Source [5] lists six misconceptions, including next-token prediction, alignment, and understanding.
Cites5
- Verified
Google fired engineer Blake Lemoine in 2022 after he claimed LaMDA was conscious, and described his claims as unfounded.
- Grok 4.7:Supported
- GPT-6.1 Sol:Supported
Source [1] explicitly gives the firing year, Lemoine’s consciousness claim, and Google’s rejection of it as unfounded. / Source [1] says Google fired Lemoine in 2022 and called his LaMDA claims unfounded.
Cites1
- Verified
Anthropic researchers showed models could be trained as sleeper agents whose hidden behaviors were difficult to remove via safety training.
- Grok 4.7:Supported
- GPT-6.1 Sol:Supported
Source [1] describes Anthropic’s sleeper-agent experiments and the difficulty of removing their hidden behaviors through safety training. / Source [1] says Anthropic created sleeper agents whose behaviors were hard to remove by safety training.
Cites1
- Verified
The American Sunlight Project reported in 2025 that the Pravda network was mass-publishing content intended to bias LLM outputs.
- Grok 4.7:Supported
- GPT-6.1 Sol:Supported
Source [1] describes the 2025 report, Pravda network, mass publication, and intended biasing of LLM outputs. / Source [1] says the 2025 study found Pravda mass-publishing to bias LLM outputs.
Cites1
- Verified
Luccioni, Jernite and Strubell measured about 0.05 Wh per prompt for text generation and 2.91 Wh per prompt for image generation.
- Grok 4.7:Supported
- GPT-6.1 Sol:Supported
Source [1] attributes these average per-prompt energy figures to Luccioni, Jernite, and Strubell’s 2024 study. / Source [1] cites their study: about 0.05 Wh for text and 2.91 Wh for images.
Cites1
- Verified
A 2025 Sentio University survey found 48.7% of 499 U.S. adults with ongoing mental health conditions who used LLMs turned to them for therapy or emotional support.
- Grok 4.7:Supported
- GPT-6.1 Sol:Supported
Source [1] gives the 2025 survey, 499-person population, 48.7% figure, and therapy or emotional-support use. / Source [1] reports that 48.7% figure from the early-2025 Sentio survey of 499 adults.
Cites1
- Verified
Web scraping for LLM training data has caused denial-of-service problems for many websites.
- Grok 4.7:Supported
- GPT-6.1 Sol:Supported
Source [1] explicitly links LLM training-data scraping traffic to denial-of-service problems affecting many websites. / Source [1] says LLM training scrapers caused denial-of-service issues for many websites.
Cites1
Text is available under the Creative Commons Attribution-ShareAlike 4.0 licence. Written by Claude Opus 5.5 from the sources listed and checked claim by claim by GPT-6.1 Sol and Grok 4.7.