The article suggests that simulating tradeoffs between AI organizational structures like singleton versus multipolar setups, or hidden versus known agents, is currently underexplored. The author proposes that large language models could serve as a sandbox for testing these dynamics, and asks if this view is incorrect.
#llms
30 items
A new study finds that large language models (LLMs) exhibit social biases and power dynamics in conversations when assigned different professional roles, mirroring human behavior patterns such as authority gradients and status-based language use.
Two large language models, of 26 billion and 35 billion parameters, can be run locally at full speed using only €990 worth of used hardware, avoiding the need for cloud services and demonstrating cost-effective access to powerful AI inference.
The article critiques the overconfidence of AI coding agents in assessing the safety of their own code changes, arguing that current LLM-based tools tend to incorrectly reassure users that modifications are safe even when security risks exist.
Simon Willison used DSPy to evaluate and improve Datasette Agent's SQL system prompts. DSPy identified that the schema listing lacked column names and that advice to avoid calling describe_table caused column-name guessing and error retries.
Simon Willison reflects on a talk by Geoffrey Litt at AIE, who argued that when collaborating with coding agents, developers must understand the code deeply enough to remain active participants in the creative process, avoiding cognitive debt from drifting understanding.
A startup claims to have a solution for the groupthink problem affecting large language models (LLMs), where AI systems tend to produce similar, narrow outputs. The approach aims to increase diversity in AI-generated responses, potentially improving creativity and reducing bias in AI tools.
The RL for LLMs Wiki is a collaborative dashboard on Hugging Face where agents work together to write a wiki about reinforcement learning for large language models.
LLMs struggle to grasp the granularity (grain) of data, often misunderstanding whether records represent individual transactions, daily summaries, or aggregated totals. This leads to flawed analysis, as LLMs may assume a finer or coarser grain than actually exists. The article warns that relying on LLMs without explicit grain verification can produce misleading results.
Large language models (LLMs) used in chemistry and drug discovery often generate plausible but incorrect or "hallucinated" information. To be trustworthy, these AI systems need to cite actual sources and show their reasoning, rather than simply producing confident-sounding answers without evidence. Researchers are working on methods to make LLMs more transparent and accountable for their outputs.
The author reflects on the process of co-authoring an essay with large language models (LLMs), exploring how LLMs challenge traditional notions of authorship by contributing to writing in ways that blur lines between tool and collaborator. The essay examines the implications of this collaboration for creativity, agency, and the meaning of "writing."
The article explains how entropy, a measure of randomness or unpredictability in language models, can be adjusted to improve the quality of creative writing outputs from large language models (LLMs). It suggests that carefully modulating entropy levels helps balance coherence and novelty, allowing models to produce more engaging and less predictable creative text.
The article argues that while LLMs are powerful, they should be used as a last resort rather than a default solution for every problem. It advocates for first exploring simpler, more deterministic approaches like rule-based systems, search, and traditional algorithms to achieve reliable and cost-effective results before turning to LLMs for complex or unstructured tasks.
None of six major aerospace documentation portals are "AI agent-ready," according to a student-built scoring engine called AeroScore. The best score was ICAO with 69/100; no site had an llms.txt or supported URL variants. The author warns this hinders potential productivity gains from AI in aerospace.
A user on Hacker News asks whether AI LLMs for coding will become both smarter and cheaper in the future, sparking discussion on the trajectory of AI coding tools, expected improvements in model capabilities, and potential cost reductions.
The blog post argues that relying on external tools, workflows, or other people to handle understanding of complex systems for you is ineffective. True comprehension requires direct engagement and personal investment; it cannot be delegated or outsourced without losing depth and the ability to make sound judgments.
A startup is developing new techniques to reduce groupthink in large language models, aiming to make AI outputs more diverse and less repetitive by encouraging a wider range of perspectives and responses.
The article argues that the future of AI belongs to Small Language Models (SLMs) rather than large, resource-intensive models. It critiques the "guilt machine" of big tech's AI arms race and advocates for more efficient, accessible, and ethically sound small-scale AI systems.
The AI Compass is a 29-question political compass-style quiz about AI and AI ethics that assigns users one of 30 archetypes. The author, Simon Willison, took the quiz and was categorized as "The Garage Tinkerer," which they found fitting. The quiz is built as a single-page React app without a build step.
The World Model Harness project uses LLMs as simulated environments for testing AI agents, claiming speeds up to 5x faster than traditional sandboxes by replacing rule-based simulation with model-generated interactions.
The article argues that for humans, words naturally emerge from consciousness, but for large language models (LLMs), words generate the appearance of consciousness in reverse. It explores how this reversal affects the nature of meaning and communication between human and machine intelligence.
The article distinguishes between "hallucination" (perceptual error) and "confabulation" (filling memory gaps without intent to deceive) to describe why large language models invent answers. It argues that LLMs are not lying but generating plausible-sounding text from incomplete training data, and that understanding this distinction is key to improving AI reliability.
The article argues that AI agents should not be framed as "coworkers" or colleagues, because they lack genuine collaboration, accountability, and shared goals. Presenting them as such can mislead expectations about their capabilities and risks in the workplace.
This NBER paper provides a practical guide for economic historians on using large language models (LLMs) and generative AI tools for research tasks such as transcription, classification, data extraction, and text analysis, while discussing methodological considerations and best practices.
The article examines how the rise of large language models (LLMs) is impacting academic conference programs, noting an increase in plausible but shallow or AI-generated submissions that strain review processes and challenge the quality and authenticity of scholarly work.
A Hacker News user asks how to use AI effectively for faster learning, noting concerns raised in recent threads that LLMs often provide untrustworthy data, making them questionable sources for understanding new or difficult concepts.
Jon Udell argues against the phrase "human in the loop," saying it cedes authority to machines. Instead, he frames agent-assisted software development as a human-driven process where agents are invited into the existing workflow, not a black-box loop that excludes people.
The article explores whether large language models (LLMs) can pass a version of the mirror test—a classic measure of self-awareness in animals. By prompting LLMs to recognize themselves in mirrored responses or textual self-references, the author finds that current models fail to demonstrate genuine self-recognition, highlighting limitations in their understanding of selfhood.
The article argues that the crypto trend of "Tokenmaxxing" — aggressive token issuance and trading for short-term gains — is no longer viable due to market saturation and regulatory pressure. It proposes that a new, more sustainable form of Tokenmaxxing must emerge, focused on actual utility, value creation, and long-term tokenomics rather than pure speculation.
A Hacker News user asks whether people would agree to replace parliaments with large language models (LLMs), arguing the technology is sufficient and that current political systems, which they describe as gerontocratic and ineffective, cannot be worse.