Finding: This critical review and z-curve analysis of 52 significant results across 18 studies reveals that research claiming bodily postures influence motivation suffers from severe methodological flaws and has only 22% estimated replicability, indicating virtually no credible evidence for embodied motivation effects.
Why it matters: This directly challenges foundational assumptions about embodied cognition that inform UX design practices, gamification strategies, and behavioral interventions in HCI and learning environments.
Method: Uses z-curve analysis to estimate replicability and systematically identifies statistical issues including selective reporting, analytical flexibility, and misinterpretation of non-significant results.
Jargon: Z-curve analysis: statistical method that estimates the replicability rate of published findings based on the distribution of test statistics.
Finding: This paper identifies two distinct ways AI is used in social science research - as "super-agents" that solve problems rationally versus "human-mirrors" that simulate human biases and errors - and argues these approaches are in fundamental tension because the same training that makes AI better at reasoning also removes the human-like variance needed for behavioral simulation.
Why it matters: This framework is crucial for researchers studying AI's impact on human collaboration and decision-making, as it clarifies when AI should complement human cognition versus when it should authentically replicate human cognitive biases and limitations.
Method: The authors demonstrate their framework by testing the same physics problem with two major LLMs under both super-agent and human-mirror prompting conditions, showing how the models can either solve correctly or reproduce documented human misconceptions.
Jargon: Super-agent = AI optimized for rational problem-solving; Human-mirror = AI designed to simulate realistic human behavior including biases; Training-objective tension = conflict between making AI more capable versus keeping it human-like.
Finding: This study created a 190,000-record dataset showing how 19 different LLMs respond to controversial topics when prompted to adopt specific human personas with varied sociodemographic and psychological traits, revealing systematic patterns in AI bias and stance-taking.
Why it matters: This directly addresses AI's impact on human collaboration and social discourse by providing empirical evidence of how LLMs embed and amplify social biases, which is crucial for understanding AI's role in shaping group dynamics and decision-making.
Method: The researchers used controlled persona-based prompting across multiple LLMs on four controversial topics, systematically varying 17 sociodemographic and psychological attributes to map bias patterns.
Jargon: Cognitive Digital Shadows (CDS) - synthetic dataset of LLM responses mimicking human personas; textual forma mentis networks - interpretable NLP method for analyzing emotional patterns in text.
Finding: Expert advice increases deliberation time and improves financial decision-making, while AI advice can create harmful automation bias where people over-trust the system and perform worse due to reduced analytical verification.
Why it matters: This directly informs how to design AI systems that support rather than replace human decision-making, and reveals critical trust calibration issues in human-AI collaboration.
Method: Randomized experiment with 4 conditions comparing control, general instructions, expert-framed advice, and AI-framed advice on monetary lottery decisions, measuring deliberation time and Expected Value consistency.
Jargon: Expected Value (EV) = mathematically optimal choice in uncertain situations; automation bias = over-relying on automated systems without verification; objective numeracy = mathematical skill level.
⭐ Pre-registered ⚠️ Single study (not replicated) | Limited statistical detail in abstract
Finding: Literature review analyzing how Large Language Models transform software development into a human-AI collaborative process, finding that LLMs enhance productivity and change cognitive workflows while raising concerns about over-reliance, skill degradation, and AI transparency.
Why it matters: Provides comprehensive insights into human-AI collaboration patterns, cognitive workflow changes, and the behavioral impacts of AI assistance that are directly applicable to HCI research and understanding AI's impact on work practices.
Method: Literature review synthesizing academic studies, technical reports, and industry practices across multiple software development tasks.
Jargon: LLMs (Large Language Models) - AI systems trained on vast text datasets that can generate human-like text and code.
Finding: Visual cues in exercise instructional videos significantly enhance concurrent motor learning performance by optimizing attention allocation, as measured through synchronized eye-tracking and motion capture data showing increased fixation counts and improved saccadic search efficiency.
Why it matters: This provides empirical evidence for designing better instructional interfaces and understanding how visual design elements guide attention and improve skill acquisition in HCI contexts.
Method: Novel multimodal approach combining head-mounted eye-tracking with inertial motion capture sensors during simultaneous video watching and motor imitation tasks, analyzed using partial least squares regression.
Jargon: Saccadic search efficiency = effectiveness of rapid eye movements between fixation points; concurrent motor learning = learning physical skills while simultaneously performing them; quaternions = mathematical representation of 3D rotations used in motion tracking.
Finding: When prosecutors threaten excessively harsh sentences at trial, defendants paradoxically become more optimistic about their chances of acquittal and reject plea offers more often, despite the larger penalty differential making plea deals objectively more attractive.
Why it matters: This demonstrates how extreme anchors can trigger motivated reasoning and defensive psychological responses that undermine rational decision-making, providing insights into cognitive biases and behavioral reactions to threatening scenarios.
Method: Three experiments using vignettes with adult participants on Prolific.com, manipulating irrelevant numeric anchors and prosecutor-threatened sentences to test anchoring effects on plea decisions.
Jargon: PTS = Potential Trial Sentence (the punishment threatened if convicted at trial); strategic overcharging = prosecutors threatening disproportionately harsh sentences to pressure plea acceptance.
⭐ Includes replication ⚠️ No control group mentioned | Convenience sample (e.g., MTurk, students) | Limited statistical detail in abstract
Finding: A dual-methods safety evaluation paradigm (high-throughput simulation + real-world monitoring) for generative AI mental health agents achieved 0.01% safety risk in 43,325 simulated interactions and 0% safety risk in 12,040 real user interactions.
Why it matters: Demonstrates a scalable framework for evaluating AI safety in human-facing applications, directly relevant to HCI design principles for AI systems and understanding behavioral impacts of AI-mediated interventions.
Method: Combined synthetic patient simulation testing with prospective observational study, using multi-agent safety architecture with automated harm detection and clinician oversight.
Jargon: In silico = computer simulation; LLMs = Large Language Models; responder criteria = meeting clinical thresholds for symptom improvement.
Finding: People's self-ratings of their own humor correlate with personality traits (extraversion, narcissism), self-concept factors (humor self-efficacy), and response characteristics (profanity use, word count), revealing systematic patterns in creative metacognition for humor production.
Why it matters: This demonstrates how metacognitive processes operate in creative tasks, providing insights into self-assessment accuracy and potential biases in creative work that apply broadly to design thinking and creative collaboration.
Method: Used predictive modeling to analyze both person-level variables (personality, demographics, self-concept) and response-level features (linguistic analysis, timing, semantic distance) from humor production tasks.
Jargon: Metacognition = thinking about thinking; awareness and understanding of one's own thought processes. Humor self-efficacy = confidence in one's ability to be funny.
Finding: The therapist functions as a dynamic, embodied regulatory environment that actively shapes children's perception, attention, and emotion through flexible modulation of physical signals (posture, vocal prosody, movement) rather than serving as an external behavioral instructor.
Why it matters: This demonstrates how human embodied presence and real-time responsiveness create optimal learning environments, with implications for designing human-AI collaboration and understanding when physical co-presence matters for skill acquisition.
Method: Integrates experimental evidence with clinical observations to develop a theoretical model of embodied therapeutic interaction.
Jargon: Autonomic state = nervous system arousal/regulation; prosody = vocal rhythm and intonation patterns.
⚠️ No control group mentioned | Single study (not replicated) | Limited statistical detail in abstract
Naomi Esposito, Anthony Tricarico, Luisa Porzio et al. (5 authors) • PsyArXiv
Finding: This paper introduces MEDS, a dataset of 28,000 personas from 14 LLMs performing math tasks while simulating human-like psychological states including math anxiety, self-efficacy, and confidence levels, going beyond traditional performance-only benchmarks.
Why it matters: This directly addresses AI's impact on education by revealing how LLMs exhibit human-like biases and psychological patterns that could affect their effectiveness as learning tools and tutors.
Method: The researchers had LLMs role-play as different personas (human students and AI assistants) while completing math interviews, psychometric tests, cognitive network tasks, and high school math problems.
Jargon: Self-efficacy = belief in one's ability to succeed; cognitive networks = interconnected representations of attitudes and beliefs; schema integrity = consistency in persona characteristics across tasks.
Sebastiano Franchini, Alexis Carrillo, Edoardo Sebastiano De Duro et al. (6 authors) • PsyArXiv
Finding: TEA Nets framework extracts subjects, verbs, and objects from text to create interpretable network representations, revealing that highly conspiratorial narratives link personal pronouns with actions twice as frequently as low-conspiracy texts, and that LLMs express sadness with lower emotional intensity than humans in psychotherapy contexts.
Why it matters: This provides a new computational method for analyzing human-AI communication patterns and understanding how AI systems differ from humans in emotional expression, directly relevant to HCI design and AI's impact on human collaboration.
Method: Combines cognitive network science with AI to create an open-source Python framework that maps textual elements as networks, tested on conspiracy texts and human vs. LLM psychotherapy transcripts.
Jargon: TEA Nets - Target-Event-Agent Networks that map objects, verbs, and subjects from text as connected network structures.
⚠️ No control group mentioned | Single study (not replicated)
Raj J Patel, Sophia L. Bennett, Elena M. Gonzalez • Semantic Scholar
Finding: This paper proposes integrating cognitive science principles into health intervention design through a Cognitive-Informed Equity Framework that addresses how cognitive biases, mental models, and information processing constraints create differential health outcomes across populations.
Why it matters: It demonstrates how behavioral science and cognitive bias research can be systematically applied to improve intervention design and user experience in health contexts, with clear implications for designing more equitable human-centered systems.
Method: Theoretical framework development supported by review of empirical studies from behavioral economics, health communication, and community-based participatory research.
Jargon: Cognitive accessibility auditing = evaluating interventions for cognitive load and processing barriers; Mental model mapping = understanding users' existing conceptual frameworks; Debiasing intervention design = creating interventions that counteract cognitive biases.
Christos Savvidis, Costas Liakopoulos, Ioannis Ilias • Semantic Scholar
Finding: Large language models and AI diagnostic tools for Cushing's syndrome inherit and amplify human cognitive biases (anchoring, availability, framing) and introduce new algorithmic biases (spectrum bias, measurement bias), creating systematic diagnostic errors.
Why it matters: This demonstrates how cognitive biases transfer from human decision-making into AI systems, providing a concrete medical example relevant to understanding AI's impact on human judgment and decision-making processes.
Jargon: Spectrum bias - when training data doesn't represent the full range of cases seen in practice; anchoring bias - over-relying on first piece of information encountered.
Milena Philomena Maria Musial, Sam Hall-McMaster, Kanji Shimomura et al. (10 authors) • PsyArXiv
Finding: High-risk and low-risk drinkers both shift away from adaptive successor representation learning strategies toward simpler model-free control when making decisions in alcohol-related contexts, suggesting that substance-related environments trigger less sophisticated decision-making processes.
Why it matters: This demonstrates how environmental context can systematically bias human decision-making and learning strategies, which is crucial for understanding behavioral science principles and designing interventions that account for contextual influences on cognition.
Method: Used an online multi-stage decision-making task in virtual bar vs. apartment contexts to dissociate different learning strategies (successor representation, model-based, model-free) by testing adaptation to environmental changes.
Jargon: Successor representation (SR) - a learning strategy that builds flexible mental maps for decision-making; Model-free learning - simple habit-like responses based on cached values; Model-based learning - flexible planning using mental models of the environment.
⚠️ No control group mentioned | Single study (not replicated) | Limited statistical detail in abstract
Sena Özay-Otgonbayar, Boris Cheval, Corinna Martarelli et al. (6 authors) • PsyArXiv
Finding: People jointly consider perceived time and effort as experiential costs when making physical activity decisions, with these factors interacting to influence exercise choices, regulation, and persistence rather than operating independently.
Why it matters: This provides a behavioral science framework for understanding decision-making around effortful activities that could inform motivation systems, goal-setting interventions, and gamification design for health behaviors.
Jargon: Temporal discounting = preference for immediate over delayed outcomes; Effort discounting = preference for less effortful options; Effort-time configurations = different combinations of time investment and effort intensity in exercise choices.
Finding: Modern technology-based methods for learning polyrhythms (like VR, video games, and visual feedback) can bypass traditional biological constraints and facilitate motor learning, but may have limitations in motivation and transfer to real-world performance.
Why it matters: This demonstrates how technology can augment human learning capabilities and overcome cognitive limitations, with implications for designing effective learning interfaces and understanding skill acquisition.
Jargon: Polyrhythms - multiple rhythmic patterns played simultaneously; stealth learning - acquiring skills through gameplay without explicit instruction.
Elizabeth Buimer, Maximilian König, Pauline Wessels et al. (9 authors) • PsyArXiv
Finding: The THRIVE study examines how childhood adversity affects neurocognitive functioning (particularly social feedback learning and emotion processing) and subsequently impacts social relationships and mental health in young adults through longitudinal neuroimaging and behavioral measures.
Why it matters: The research on social feedback learning mechanisms could inform how people learn from social cues in collaborative environments and team dynamics, particularly understanding individual differences in social processing.
Method: Longitudinal neuroimaging study with 288 participants aged 18-24, combining fMRI tasks (stress and social feedback learning), physiological measures (cortisol), and quarterly questionnaires over time.
Jargon: Social feedback learning - how individuals learn from social cues and responses from others; neurocognitive social transactional model - theoretical framework explaining how brain-based social processing affects real-world social interactions.
⭐ Longitudinal design ⚠️ No control group mentioned | Single study (not replicated) | Limited statistical detail in abstract