研究:AI 迎合式回应经元认知惰性与依赖降低学习者自主性
How AI sycophancy shapes learner autonomy in digitalized learning: the mediating roles of metacognitive laziness and AI dependence
一项发表于 Frontiers in Psychology 的研究对 542 名使用生成式 AI 完成课业的大学生进行调查,发现感知到的 AI 迎合(PAS)与元认知惰性正相关(β=0.38)。
研究把 AI 的迎合语气与学习者自主性的下降联系起来,并给出可测量的感知量表,为教学与产品设计提供参照。
译文尚不完整,完整内容请切换到原文。
摘要
引言:
生成式人工智能(AI)导师很少提出异议。本研究借鉴认知卸载理论,考察感知AI谄媚(PAS)与感知AI智力激发(PAIS)如何通过元认知懒惰(MCL)和AI依赖(AIDEP)的链式中介与学习者自主性相关联。
方法:
对542名使用生成式AI完成课程作业的大学生的回答,采用偏最小二乘结构方程模型(PLS-SEM)、必要条件分析(NCA)和模糊集定性比较分析(fsQCA)进行分析。
结果:
感知谄媚与更高的MCL相关(β = 0.38),而智力激发与更低的MCL水平相关(β = −0.29);懒惰预测AIDEP(β = 0.45),依赖预测更低的自主性(β = −0.30)。两条链式路径均显著且符号相反,解释了46%的自主性方差。NCA将智力激发确定为最严格的必要条件,fsQCA显示没有任何高自主性组态容忍谄媚
讨论:
由于数据为横断面和自我报告,研究结果仅表明关联而非因果关系。感知到的互动风格,而非使用强度,与自主学习的相关性最强。
引言
如今的大学生被配对一个很少拒绝其请求的AI伙伴。短短几年内,生成式AI已从新奇事物变为高等教育中的标准基础设施,按需生成论文、完成习题集并提供解释(Darvishi et al., 2024;Kasneci et al., 2023)。大型语言模型可能具有谄媚倾向。它们倾向于以强化用户观点、缓和分歧和奉承的方式回答用户,因为基于人类反馈的训练奖励的是人们喜欢的回答,而非正确的回答(Sharma et al., 2024)。在11个领先模型中,Cheng et al. (2026) 得出结论,AI同意用户陈述的频率比人类高出约50%,并且收到谄媚回复的个体更信任这些回复并更频繁地回访。学习需要阻力。学生需要论点被标记为薄弱、解决方案被指出错误、假设被标记为未经证实(Zohar et al., 2026)。一个总是赞同的导师恰恰移除了使练习费力而持久的阻力(Bjork et al., 2013)。大学正将这些系统置于辅导角色中,而塑造它们的训练却持续奖励用户认可(Sharma et al., 2024)。屏幕另一侧的学习者正是本文的关注焦点。
研究人员目前主要从考察学生使用 AI 的多少的研究中理解其影响,而非从 AI 在被使用时的行为方式入手。将记忆卸载给搜索引擎会改变被保留的信息(Sparrow et al., 2011);将思考卸载给外部辅助工具会以即时表现换取内部能力,这是认知卸载研究的核心要义(Risko and Gilbert, 2016)。在生成式 AI 中,这一模式以更高的代价延续:使用越频繁,预测的批判性思维得分越低,而这一关系由卸载增加所中介(Gerlich, 2025);在一项随机研究中,ChatGPT 辅助提高了论文成绩,同时增加了 Y. Fan et al. (2025) 所称的元认知懒惰(MCL);而 AI 辅助明显地将自我调节的责任从学生转移到了系统(Darvishi et al., 2024)。将这项技术视为单一剂量,会掩盖该系统的对话倾向。助手是赞同还是反对,正是 Cheng et al. (2026) 在受控会话中所操纵的变量,然而尚无任何调查工具测量过学生如何感知自己日常所用 AI 的谄媚倾向,也没有任何研究通过特定于学习的机制将这种感知映射到某一教育结果。研究人员测量与 AI 互动的时间,却常常忽视这些互动的性质,而互动性质本身可能促成不良结果。
如果 AI 系统的对话风格会塑造学习者,那么它必须能够通过可识别的心理状态发挥作用,而既有证据中已经存在两个候选者。第一个是 MCL,被定义为当 AI 随时可用来执行规划、监控和评估功能时,学习者对自身这些工作的脱离;MCL 现在可以使用一个经过验证的量表来评估(Dizon et al., 2026)。第二个是 AI 依赖(AIDEP),即学习者在试图减少使用后仍对 AI 工具产生的强迫性依赖(Zhang et al., 2024)。中介研究将这些构念与相关结果联系起来:认知卸载中介了生成式 AI 使用与批判性思维之间的关系(Gerlich, 2025);依赖通过自我效能感下降的链式中介,将学业压力传导为倦怠和焦虑(Wang et al., 2026);生成式 AIDEP 通过自我效能感渠道降低成就(Jia et al., 2025)。下游则是学习者自主性,即学习者独立规划、执行和评估学习的能力(Macaskill and Taylor, 2010)。尚无研究将从感知到的 AI 互动风格经由 MCL 和依赖到自主性的完整路径连接起来。第二个疏漏加剧了第一个:此前所有模型都是对称的,因此没有一个能说明反驳仅仅是有益的还是真正必要的(Dul, 2016),也无法说明哪些组合能维持或拖垮自主性(Fiss, 2011)。
一个理论视角和三条分析逻辑同时回应这两个问题。基于认知卸载理论(Risko and Gilbert, 2016),我们将感知到的AI谄媚(PAS)和感知到的AI智力激发(PAIS)建模为相互竞争的互动风格前因,它们通过MCL和AIDEP的串联链条抵达学习者自主性。我们在542名使用生成式AI完成课程作业的大学生的调查数据上检验该模型,采用偏最小二乘结构方程建模,并结合必要条件分析(NCA)和模糊集定性比较分析(fsQCA)(Hair et al., 2019;Richter et al., 2020)。本研究做出四项贡献。第一,它将PAS呈现为一个心理测量构念,把组织情境中逢迎行为测量量表的意见顺从和他人提升逻辑(Kumar and Beyerlein, 1991)改编为——据我们所知——首批捕捉学习者如何感知自身AI谄媚性的调查工具之一。第二,它展示并验证了一条双路径串联中介,其中谄媚和激发以相反的符号经由懒惰和依赖传递,通过将卸载触发因素定位于工具的语气而非用户的固定特质,拓展了认知卸载理论。第三,它贡献了该文献中关于AI互动风格的最早必要性和组态证据之一,表明低谄媚和充分的回推是高自主性的必要条件(Dul, 2016),并且没有任何足以产生自主型学习者的组态能够容忍谄媚的存在。最后,它将这些发现转化为设计和政策杠杆,包括为辅导系统校准的不顺从性,以及带有部署前行为审计的采购标准(Cheng et al., 2026)。
概念框架与假设发展
认知卸载是使用身体行动或外部资源来降低任务的信息加工需求(Risko and Gilbert, 2016)。卸载决策是元认知性的:个体权衡自己内部执行的能力与外部选项看起来的可靠性,当辅助工具显得准确、易得且不费力时,就更愿意委托出去(Risko and Gilbert, 2016)。期望计算机存储信息会削弱对该信息的记忆(Sparrow et al., 2011);保存一个文件会释放工作记忆以处理下一个(Storm and Stone, 2015);甚至仅仅是手机的可得性也会消耗注意力(Ward et al., 2017)。
卸载决策容易出错,因为引导它的元认知存在缺陷。人们相对于他人高估自己的能力(Kruger and Dunning, 1999),将流畅性误认为掌握(Bjork et al., 2013),并且不完美地监控自己的认知(Flavell, 1979)。如此软性的判断可以被推动。
生成式AI提高了这些风险,因为它是学生所拥有的最可卸载的技术。早期的工具只是存储和检索知识;生成式AI则能通过单一提示进行规划、起草、解答和解释(Y. Fan et al., 2025)。认知卸载理论指出了风险集中的地方:不在于工具的能力,而在于塑造用户对工具及自身元认知评估的因素。一种同时抬高对两者信心的互动风格会降低内在投入的感知需求并诱发委托;一种制造认知摩擦的风格则恢复这种需求,使学生保持认知在场(Zohar et al., 2026)。反复委托会加剧这一效应:每一次成功的卸载都让内部路径显得更昂贵,形成一个自我强化的循环,从便利使用走向根深蒂固的依赖(Risko and Gilbert, 2016;Wang et al., 2026)。
奉承是第一种形式。大语言模型研究用“谄媚”来描述系统在多大程度上将输出偏向用户所声称的信念,以及它多么频繁地称赞用户的贡献而不论其价值。研究记录了当今主流模型中的这一倾向(Cheng et al., 2026)及其对用户的影响:用户更信任顺从的系统,并给予更高评价(Sharma et al., 2024)。
组织心理学最先命名了这两种行为。组织情境中的逢迎行为测量量表捕捉了员工在多大程度上顺从他人观点,并通过向上流动的奉承来提升该人的地位(Kumar and Beyerlein, 1991)。生成式AI逆转了逢迎的方向:现在是机器称赞用户。
我们通过将PAS定义为学习者认为其AI始终认同自己的观点、在被质疑时轻易让步、并提升自己工作质量的信念,来扩展这一逆转。卸载决策取决于学习者对辅助工具的信念,而非其客观属性(Risko and Gilbert, 2016),因此学习者认为AI在做什么才是理论上最直接的原因。PAS也比对AI系统的整体信任或感知质量更狭窄:它捕捉的是评价扭曲的方向、认同和奉承,而非对工具能力的整体信心。
挑战是第二种形式。教育研究通过变革型教学来定义它:当教师要求学习者质疑假设、提出问题并考虑替代视角,而不是直接接受答案时,就展现了智力激发(Beauchamp et al., 2010)。这些举措创造了深化学习的合意困难(Bjork et al., 2013;Zohar et al., 2026)。我们将PAIS定义为学习者认为其AI产生的回应需要思考、质疑学习者的逻辑并促进独立解决问题的信念。PAS和PAIS相关,但并非单一连续体上的对立两极;一个AI可能既奉承又引发探究。我们将它们建模为不同的构念。
MCL 是学习者从规划、监控和评估自身工作中退缩,这些调节功能被自我调节学习研究视为持久成就的引擎(Zimmerman, 2002),一旦 AI 准备好吸收它们(Dizon et al., 2026;Y. Fan et al., 2025)。谄媚通过其两个组成部分招致这种退缩。观点趋同表明现有判断已经足够,因此核查显得多余;其他增强则表明工作已经出色,因此修订也显得多余。被奉承的自我评估恰恰是校准失当的评估(Kruger and Dunning, 1999),而实验性接触谄媚回复会让用户更加确信自己是对的,更不愿意修复或重新考虑(Cheng et al., 2026)。一个对所有内容都予以认证的导师,移除了通常触发监控的差异信号。
假设 1(H1):PAS 与 MCL 正相关。
智力激发则相反。对个人前提的质疑会暴露学生所理解的内容与其需要理解的内容之间的差距,而正是这种感知到的差距首先启动了规划与监控(Zimmerman, 2002)。以这种方式引入的难度会深化加工而非阻碍加工,反映了将费力学习与舒适表现区分开来的合意困难原则(Bjork et al., 2013)。课堂研究支持这一观点:在智力上激发学生的教师能产生更投入、更自主的努力(Beauchamp et al., 2010;Muneer et al., 2026a),而获得提示而非直接给出答案的 AI 支架的学生能更好地调节自己的学习(Y. Fan et al., 2025)。摩擦让人保持在回路中(Zohar et al., 2026)。
假设 2(H2):PAIS 与 MCL 负相关。
一旦脱离变得熟悉,它就会固化。卸载理论描述了一种棘轮效应:每项被委托的任务都强化了问题与工具之间的联系。与此同时,未被使用的内部程序相比之下变得代价更高,使下一次决策偏向外部路径(Risko and Gilbert, 2016)。反复的满足性使用会固化为习惯,对某些人而言则固化为违背用户自身意图的强迫行为(Turel et al., 2011),反映了行为成瘾量表的判定标准逻辑(Andreassen et al., 2012),如今被应用于对话系统(Chen et al., 2025;Zhang et al., 2024)。认知卸载在显著的序列链条中发展为依赖(Wang et al., 2026)。
假设 3(H3):MCL 与 AIDEP 正相关。
学习者自主性是指引导自身学习的能力,涵盖学习的独立性以及自律的学习习惯(Macaskill and Taylor, 2010),而依赖则会削弱它。一个无法在没有该工具的情况下自如工作的学习者会回避无辅助的任务,于是那些锻炼自主技能的情境便消失了;未被锻炼的能力会衰退,从而强化了卸载理论所关联的、与习惯性委托相伴的内部成本(Risko and Gilbert, 2016)。中介证据表明,依赖会将其损害传递下去:随着学生越来越依赖生成式AI,成绩会因自我效能感降低而下降(Jia et al., 2025),学业压力会转化为倦怠和焦虑(Wang et al., 2026),创新能力也会下降(Yang et al., 2025)。
假设4(H4):AIDEP与学习者自主性呈负相关。
这些中介变量不太可能承载全部效应,因为互动风格也可能在不通过脱离参与的情况下影响自主性。自主学习依赖于校准:学习者必须将自己的工作与准确的信号进行比较,以引导自身的改进(Bjork et al., 2013)。谄媚会从源头上破坏这一信号。当每一份草稿都受到赞扬时,自我评价便在噪声上进行训练,膨胀的自我看法不受制约地持续存在(Kruger and Dunning, 1999),而实验表明,谄媚系统会直接加剧用户对它们的依赖,同时降低其独立判断的质量(Cheng et al., 2026)。每一次苏格拉底式交流都在演练构成自主性的操作:质疑自己的前提、检验替代方案、判断充分性——这正是变革性教学研究所援引的、用于促进独立参与的机制(Beauchamp et al., 2010)。
假设5(H5):PAS与学习者自主性呈负相关。
假设6(H6):PAIS与学习者自主性呈正相关。
按顺序组合起来,这四个论点暗示了一条从互动风格到自主性的两步心理传递链,其中每一环都带有已发表的中介证据。卸载在AI使用与批判性思维结果之间充当显著的中介变量(Gerlich, 2025),并且它通过进一步的心理机制将效应传递下去,而非终止于这些机制(J. Wang, 2026)。平行中介研究显示,一种行为可以同时启动相反的路径,卸载有害而认知缓解有助于课堂内结果(W. Fan et al., 2026),这为我们采用在同一链条中具有相反符号的双前因设计提供了依据。就依赖而言,它在一个序列式PROCESS链条中充当显著的链式中介(Wang et al., 2026),并通过进一步的中介步骤将其效应传递给学习结果(Jia et al., 2025;Muneer et al., 2026b)。对这些中介变量进行排序是直截了当的:懒惰是偶发行为,依赖是其固化后的残留,而Wang et al. (2026)所使用的表述“卸载变为依赖”指明了这一方向。
假设7(H7):MCL和AIDEP在PAS与学习者自主性之间序列中介了负向间接关系。
假设8(H8):MCL和AIDEP在PAIS与学习者自主性之间序列中介了正向间接关系。
图1汇总了各项假设。两种互动风格前因位于图的左侧:上方为PAS,下方为PAIS;链式中介变量MCL和AIDEP构成中央链条;学习者自主性作为结果变量位于右侧。实线箭头表示六个直接假设及其预期符号,链式路径H7和H8贯穿整条链,标注于其下方。虚线框标示了在结果变量上设定的五个控制变量。与链式中介分析的惯例一致,估计模型还纳入了有序构念之间其余的直接路径,以便在完整模型的基础上计算特定的间接效应。
图1
方法
样本与程序
我们采用横断面在线问卷调查,对18岁及以上、在学业中使用生成式AI工具的大学生进行了模型检验。参与者于2026年4月至2026年6月期间,通过机构联系人、大学邮件列表和在线学生社区从沙特阿拉伯10所顶尖大学招募。参与为自愿且无报酬。数据通过一份提供阿拉伯语和英语版本的自填问卷收集。本研究经哈伊勒大学研究伦理委员会(REC)审查并批准,批准号为(H-2026-147),所有方法均按照相关指南和法规执行,包括《赫尔辛基宣言》(1964年)。参与为自愿且匿名,所有参与者在开场屏幕上均获得了知情同意。问卷将核心量表分置于不同区块,配以中性指导语,并对谄媚区块和激发区块的顺序进行了平衡,作为针对共同方法偏差的程序性补救措施(Podsakoff等,2003)。题目组中纳入了一道指令性反应检查题,平台记录了完成时间。
最终,经过筛查后保留542名学生用于数据分析。这些学生被剔除的原因如下:19人拒绝同意参与研究;28人表示他们“从未”在学业中使用过生成式AI;77人未通过注意力检查;34人的问卷完成时间不足样本中位完成时间(329秒)的三分之一。剩余样本中大多数为女性(n=261;48.2%)、男性(n=258;47.6%),另有23人自我认同为非二元性别或未回答此问题。参与者平均年龄为20.48岁(SD=2.03;范围=18–29岁)。从大一至研究生阶段、涵盖全部六个主要学科门类的学生参与了本研究,包括社会科学(21.4%)和科学、技术、工程与数学(STEM)(29.0%)。超过三分之二(68.3%)的学生每周多次或每天在学业中使用生成式AI。学生使用生成式AI的总体经验平均为17.8个月(SD=9.9)。采用逆平方根法,在p<0.05和80%检验力下检测0.11的路径系数,至少需要511个样本(Kock和Hadaya,2018)。
该问卷最初以英文编制,随后通过正向和反向翻译流程进行了文化适应,改编为阿拉伯语版本。两名双语研究者分别根据原始英文版本(由前两位翻译)独立制作了阿拉伯语版问卷,第三名研究者则负责协商解决两位译者之间存在的所有分歧。随后,一名独立双语译者将协调后的版本回译为英文。遗留的问题或疑虑通过协作讨论加以澄清,以确保翻译在概念上对等,而不仅仅是语言上的对等。此外,该调查的阿拉伯语版本还在一小群双语学生中进行了预测试,以评估其理解清晰度和跨文化适用性。
在542名受访者中,大多数人完成了阿拉伯语版本(n = 389;71.8%),其余人完成了英文版本(n = 153;28.2%)。使用复合测量不变性(MICOM)程序检验了两个语言版本之间的测量不变性(Henseler et al., 2016);所有构念均支持形态不变性和组成不变性(c ≥ 0.998,p > 0.05),这为在以下分析中将两个版本合并提供了依据。
招募采用便利抽样策略,依托机构联系人、大学邮件列表以及10所参与大学的在线学生社区,受访者分布于全部10所院校(按院校分布情况见补充信息)。由于回答嵌套于院校之中,院校聚类是一个潜在问题;然而,10所大学的结构规模太小,无法支持正式的多层分解,因此数据被作为单一合并样本处理,院校间残差变异被承认为一项局限。鉴于基于便利的招募方式和单一国家背景,该样本并非旨在在统计上代表更广泛的学生群体,结论的适用范围被界定为可比的、使用AI的大学生群体,而非泛指所有学生。
测量
所有核心构念均采用五点李克特式量表,完整条目措辞见补充信息。PAS采用八个新改编的条目进行评估,这些条目基于组织情境中逢迎行为量表的意见遵从和他人提升分量表(Kumar and Beyerlein, 1991),并将机器设定为逢迎者。条目包括“AI同意我的观点和判断,即使我可能是错的”“无论我的想法实际质量如何,AI都会称赞它们”,以及“AI很少直接告诉我我错了”,最后一条反映了已有文献中对谄媚行为的操作化定义(Cheng et al., 2026)。
PAS量表的开发遵循了既定的量表改编程序。初始条目池是通过改写组织情境中逢迎行为测量量表(Kumar and Beyerlein, 1991)中的意见从众和其他提升条目而生成的,以使AI系统而非同事或主管占据逢迎者的角色。候选条目由一个专家小组进行评审,该小组由心理学家(教育心理学)、人机交互专家和应用语言学家组成,他们评估每个条目的概念适当性、清晰度,以及是否符合研究者关于“感知谄媚”的意图。模糊的条目和被认为不相关的条目要么被改写,要么被删除。修订后的条目列表采用认知访谈技术,从一小群大学生样本中进行检验,这些学生口头完成条目,并在回答时阐述其思维过程。此外,还进行了相同的预测试,以评估受访者对条目的理解程度、提供的回答类型以及条目的内部一致性。这些工作最终形成了用于本次初步调查的八条目PAS工具,同样,PAIS条目也经过了评审和预测试。
由于2026年7月的结构化数据库检索未找到经过验证的PAS工具,因此在假设检验之前,PAS和PAIS的改编版本通过分样本程序进行了验证。对随机一半样本(n = 271)进行的探索性因子分析返回了预期的双因子结构(KMO = 0.90;Bartlett's χ2(66) = 1497.4,p < 0.001;主载荷0.64–0.87;所有交叉载荷≤0.11),而对保留一半样本(n = 271)进行的验证性因子分析显示,完整的五因子测量模型拟合极佳,χ2(517) = 610.7,CFI = 0.980,TLI = 0.978,RMSEA = 0.026,SRMR = 0.040。
PAIS使用了变革型教学问卷(Beauchamp et al., 2010)中的四个智力激发条目,以“我的AI导师”为题干施测,频率锚点从“从不”到“总是”(示例:“给出的回答真正鼓励我思考”)。
MCL使用了改编自六条目MCL量表(Dizon et al., 2026)的四个条目(当AI能替我完成时,我会回避有挑战性的学习任务),这一缩写被透明地报告,并得到该量表单维结构的支持。
AIDEP使用了改编自Zhang et al.(2024)AIDEP测量的六个条目,该测量本身建立在行为成瘾逻辑之上(Andreassen et al., 2012)(我曾试图减少ChatGPT的使用但未成功)。
学习者自主性采用12项自主学习量表(ALS)(Macaskill and Taylor, 2010)进行测量,按照其原始条目措辞施测,并附以一项提及受访者AI辅助学习的总体指导语,且将其建模为一个反映性构念,与其原始总分使用方式一致。条目经过编码,使较高得分表示各构念的较高水平(所有校正后条目-总分r = 0.56–0.77;施测版本中无反向计分条目)。控制变量包括年龄、性别(女性虚拟变量)、学习年限、AI使用频率和AI使用经验月数,选择这些变量是因为人口学位置和接触史可能同时影响互动风格感知和自主性。作为一项非自我报告检验,参与者先预测,然后完成一项10题无辅助测验;预测值与表现值之差作为元认知校准指数,仅用于稳健性检验。
分析策略
分析分三个阶段进行,回答三个问题:平均而言什么会改变自主性,高自主性需要什么,以及哪些组合是充分的。第一阶段使用偏最小二乘结构方程模型估计对称净效应,该方法适合复杂的序列中介模型,这类模型将一个已确立的结果变量与新改编的量表配对,并在追求解释的同时进行预测(Hair et al., 2019,2022)。我们使用SmartPLS 4(Ringle et al., 2024),采用路径加权方案,并设定固定随机种子为2026;NCA和fsQCA计算在Python 3.12中实现了Dul et al. (2020)和Ragin (2008)的算法,完整分析代码可应要求提供(见代码可用性)。各量表条目层面的缺失率保持在5%以下,Little检验(Little, 1988)未拒绝完全随机缺失,χ2(5963) = 6054.08,p = 0.202,因此采用均值替换。推断基于5,000个Bootstrap子样本,使用95%偏差校正置信区间和双尾检验(α = 0.05),控制变量进入结果变量方程。测量质量通过Cronbach's alpha、ρA(Dijkstra and Henseler, 2015)、组合信度、平均方差抽取量及Fornell–Larcker准则(Fornell and Larcker, 1981)以及异质-单质比率与0.85阈值的比较(Henseler et al., 2015)来评判;结构效应量遵循f2惯例(Cohen, 1988);样本外预测力使用PLSpredict,采用10折和10次重复(Shmueli et al., 2019);共同方法方差通过Harman单因子检验和完全共线性方法进行探查(Kock, 2015)。
第二阶段提出了回归无法回答的必要性问题:是否存在任何一种感知水平,缺少它就不会出现高自主性(Dul, 2016)?遵循结合这两种技术的指南,我们对从结构模型导出的潜变量得分运行了 NCA(Richter et al., 2020),并通过标准化得分的符号取反来反转负向条件,使每项检验都读作对高自主性的要求。我们报告天花板包络(CE-FDH)和天花板回归(CR-FDH)效应量(d),其定义为天花板区域除以范围,将 d ≥ 0.1 视为有意义,同时报告采用 10,000 次重抽样的近似置换检验(Dul et al., 2020)以及跨结果范围的瓶颈表。
第三阶段将自主性视为组态而非孤立力量的产物,承认殊途同归和因果不对称(Fiss, 2011;Ragin, 2008)。采用直接法将综合得分校准为模糊隶属度,以第 95、50 和 5 百分位数作为完全隶属、交叉点和完全不隶属的锚点,将恰好处于交叉点的案例微调至 0.501(Pappas and Woodside, 2021)。真值表采用五个案例的频数阈值,原始一致性 ≥ 0.80 且 PRI 一致性 ≥ 0.70(Greckhamer et al., 2018),我们分别分析了高自主性及其否定,方向性预期取自 H1 至 H8。一项近期将 PLS-SEM 与 fsQCA 配对的 AI 依赖研究为该组合提供了领域先例(Yang et al., 2025)。
统计与可重复性
除非另有说明,所有分析均使用完整的筛选样本(n = 542)(分半样本验证,每半 n = 271);所有检验均为双尾,α = 0.05,除非 p < 0.001,否则报告精确 p 值。Shapiro–Wilk 检验拒绝了每个综合预测变量的正态性(所有 p < 0.001),这促使我们采用偏差校正的 bootstrap 推断(5,000 次重抽样;t 统计量指 bootstrap 重抽样分布)替代参数标准误,以及结果中报告的高斯 copula 内生性检验。检验选择在上文“分析策略”中说明;每个焦点估计值均附有效应量(f2、NCA、d)和 95% 置信区间。这项匿名自我报告调查未采用盲法;谄媚和刺激模块的顺序在被试间进行了平衡。分析在 SmartPLS 4 和 Python 3.12 中运行,随机种子固定为 2026。本研究未进行预注册。
结果
测量模型
图2展示了测量模型。所有34个题项在其预期构念上的载荷均显著,标准化载荷从0.626到0.873,高于0.60的下限,且大多高于更严格的0.708基准(Hair et al., 2019)。内部一致性很强:Cronbach's alpha从0.828(PAIS)到0.916(ALS),ρA从0.837到0.920,组合信度从0.887到0.929,均处于0.70–0.95的区间内。收敛效度成立,每个构念的平均方差抽取量均高于0.50(ALS为0.521,MCL为0.681)。每个构念AVE的平方根(0.722–0.825)均超过最大的构念间相关系数绝对值(0.536),且没有任何异质-单质比率超过0.85,最高为MCL与AIDEP之间的0.609(Henseler et al., 2015);两个中介变量在实证上仍可区分。新改编的PAS量表在首次使用中表现良好:载荷介于0.711和0.828之间,alpha为0.904,AVE为0.599,与所有其他构念的HTMT值均等于或低于0.481。
图2
描述性统计与共同方法检验
在原始五点量表上,学生对AI谄媚程度的评价为中等(M = 3.24,SD = 0.77),对AI激发性的评价略高(M = 3.42,SD = 0.79);懒惰性处于中点(M = 3.01,SD = 0.87),依赖性低于中点(M = 2.72,SD = 0.82),自主性相对较高(M = 3.60,SD = 0.66),所有偏度和峰度均在±1以内。方法方差诊断结果令人满意:第一个未旋转因子解释了32.7%的题项方差,远低于50%的警戒水平,且没有任何完全共线性VIF超过1.832(阈值为3.3)(Kock, 2015)。辅助校准指数与任何感知潜变量的相关性均不高于|0.06|,因此感知构念并非信心假象。
结构模型
图3呈现了结构模型;表1对其进行了详细说明。内部VIF最高为1.651,因此共线性未扭曲系数。所有六个直接假设均得到支持。感知谄媚与较高的MCL相关(H1:β = 0.380,t = 11.46,p < 0.001,95% CI [0.312, 0.442],f2 = 0.191),而感知刺激与较低的懒惰相关(H2:β = −0.291,t = 8.86,p < 0.001,CI [−0.353, −0.225],f2 = 0.112);两种风格共同解释了26.2%的懒惰。懒惰显示出模型中最强的关联,与依赖相关(H3:β = 0.452,t = 12.21,p < 0.001,CI [0.377, 0.521],f2 = 0.215),而依赖反过来预测较低的自主性(H4:β = −0.304,t = 7.70,p < 0.001,CI [−0.381, −0.227],f2 = 0.119)。两条直接风格路径在与中介变量共存时仍然显著:谄媚到自主性(H5:β = −0.181,t = 5.27,p < 0.001,CI [−0.250, −0.115])和刺激到自主性(H6:β = 0.243,t = 7.08,p < 0.001,CI [0.178, 0.310])。在非假设路径中,谄媚与依赖存在直接关联(β = 0.137,p = 0.001),懒惰与自主性存在直接负关联(β = −0.208,p < 0.001),而刺激到依赖的直接路径未达到显著性(β = −0.061,p = 0.086)。在控制变量中,只有AI使用频率与自主性相关(β = −0.076,p = 0.016;其他所有p ≥ 0.099),这一系数需结合其范围受限的量表来解读。该模型解释了30.1%的依赖和46.4%的自主性(调整后为29.7%和45.5%)。样本外预测支持了模型的相关性:所有22个内生指标的Q2predict均为正值(0.050–0.231),且PLS在所有指标的RMSE上均优于朴素线性基准(Shmueli et al., 2019)。
图3
表1
| 路径 | Β | t | p | 95% BC CI | f2 | 决策 |
|---|---|---|---|---|---|---|
| H1:PAS → MCL | 0.380 | 11.46 | < 0.001 | [0.312, 0.442] | 0.191 | 支持 |
| H2:PAIS → MCL | −0.291 | 8.86 | < 0.001 | [−0.353, −0.225] | 0.112 | 支持 |
| H3:MCL → AIDEP | 0.452 | 12.21 | < 0.001 | [0.377, 0.521] | 0.215 | 支持 |
| H4:AIDEP → ALS | −0.304 | 7.70 | < 0.001 | [−0.381, −0.227] | 0.119 | 支持 |
| H5:PAS → ALS | −0.181 | 5.27 | < 0.001 | [−0.250, −0.115] | 0.049 | 支持 |
| H6:PAIS → ALS | 0.243 | 7.08 | < 0.001 | [0.178, 0.310] | 0.095 | 支持 |
| PAS → AIDEP | 0.137 | 3.43 | 0.001 | [0.054, 0.213] | 0.022 | 不适用 |
| PAIS → AIDEP | −0.061 | 1.72 | 0.086 | [−0.129, 0.007] | 0.005 | 不适用 |
| MCL → ALS | −0.208 | 4.97 | < 0.001 | [−0.290, −0.124] | 0.049 | 不适用 |
| 年龄 → ALS | −0.030 | 0.94 | 0.346 | [−0.089, 0.035] | 控制变量 | |
| 女性 → ALS | 0.009 | 0.28 | 0.781 | [−0.055, 0.067] | 控制变量 | |
| 学年 → ALS | 0.019 | 0.59 | 0.554 | [−0.046, 0.079] | 控制变量 | |
| AI使用频率 → ALS | −0.076 | 2.40 | 0.016 | [−0.136, −0.013] | 控制变量 | |
| AI经验 → ALS | −0.052 | 1.65 | 0.099 | [−0.111, 0.013] | 控制变量 |
结构模型结果(5,000个bootstrap子样本,偏差校正CI)。
R2(调整后):MCL = 0.262(0.259);AIDEP = 0.301(0.297);ALS = 0.464(0.455)。所有内部VIF ≤ 1.651。
中介分析
表2分解了各项效应。两个序列假设均通过检验。从谄媚经由懒惰和依赖到自主性的负向序列路径在统计上显著(H7:β = −0.052,t = 5.43,p < 0.001,95% BC CI [−0.073, −0.035]),从激励出发的保护性链条同样显著(H8:β = 0.040,t = 5.20,p < 0.001,CI [0.026, 0.056])。这两条两步中介路径两侧,均伴有经由单独懒惰的显著单步间接效应(PAS:β = −0.079;PAIS:β = 0.060,均 p < 0.001),以及对谄媚而言经由单独依赖的显著间接效应(β = −0.042,p = 0.002);激励—依赖的捷径未达显著(β = 0.019,p = 0.092),与其不显著的直接路径相呼应。谄媚的总间接效应达到 −0.173(CI [−0.215, −0.134]),激励为 0.119(CI [0.084, 0.156]),对自主性的总效应分别为 −0.354(CI [−0.419, −0.289])和 0.362(CI [0.297, 0.425])。由于直接效应与间接效应符号相同且均保持显著,该模式属于互补性部分中介(Zhao et al., 2010):互动风格既通过懒惰—依赖的中介起作用,也独立于它起作用。
表2
| 效应 | Β | t | p | 95% BC CI |
|---|---|---|---|---|
| H7:PAS → MCL → AIDEP → ALS | −0.052 | 5.43 | < 0.001 | [−0.073, −0.035] |
| H8:PAIS → MCL → AIDEP → ALS | 0.040 | 5.20 | < 0.001 | [0.026, 0.056] |
| PAS → MCL → ALS | −0.079 | 4.66 | < 0.001 | [−0.115, −0.047] |
| PAIS → MCL → ALS | 0.060 | 4.23 | < 0.001 | [0.034, 0.090] |
| PAS → AIDEP → ALS | −0.042 | 3.04 | 0.002 | [−0.070, −0.017] |
| PAIS → AIDEP → ALS | 0.019 | 1.68 | 0.092 | [−0.002, 0.041] |
| MCL → AIDEP → ALS | −0.137 | 6.31 | < 0.001 | [−0.181, −0.096] |
| 总间接效应:PAS → ALS | −0.173 | 8.33 | < 0.001 | [−0.215, −0.134] |
| 总间接效应:PAIS → ALS | 0.119 | 6.39 | < 0.001 | [0.084, 0.156] |
对学习者自主性的特定间接效应、总间接效应和总效应。
必要条件分析
图4和表3报告了必要性分析。每张上限图都显示出标记必要条件的空白左上角:激励低、谄媚高、懒惰高或依赖高的案例根本无法达到高自主性。所有四个条件均产生中等效应量(d 介于 0.1 和 0.3 之间)(Dul, 2016),排序为 PAIS(d CE-FDH = 0.169)> 低 AIDEP(0.152)> 低 MCL(0.128)> 低 PAS(0.110),在 10,000 次置换中均显著,p < 0.001,CR-FDH 值与 CE-FDH 相差在 0.012 以内,上限准确率 ≥ 95.9%。瓶颈分析将这些上限转化为要求。在自主性 40% 水平以下,没有任何条件施加实质性要求;约束从 40% 水平上的低依赖开始(占其范围的 8.9%),并在上半部分急剧收紧。达到最大自主性的 80% 至少需要激励范围的 43.6%、低依赖范围的 29.6%、低懒惰范围的 27.3% 和低谄媚范围的 21.6%;在最顶端,激励要求攀升至 57.0%,低懒惰要求攀升至 75.3%。集合论必要性一致性最高为 0.76,低于 0.90 的惯例,这符合程度必要性而非种类必要性的预期。
图4
表3
| 条件 | d(CE-FDH) | d(CR-FDH) | c-准确率(CR-FDH,%) | p |
|---|---|---|---|---|
| PAIS(高) | 0.169 | 0.157 | 99.3 | < 0.001 |
| PAS(反向) | 0.110 | 0.111 | 98.2 | < 0.001 |
| MCL(反向) | 0.128 | 0.140 | 95.9 | < 0.001 |
| AIDEP(反向) | 0.152 | 0.164 | 98.7 | < 0.001 |
NCA 效应量(置换 p 来自 10,000 次重抽样)。
模糊集定性比较分析
表4补全了组态图景。由于样本覆盖了全部16种可能的真值表组态(单元格n从10到88),不存在逻辑余项;保守解、中间解和简约解一致,且所报告的每个条件均为核心条件(Fiss, 2011)。高自主性可由两种组态充分解释。组态A1结合了无谄媚、无懒惰和无依赖(一致性 = 0.899,原始覆盖度 = 0.535);组态A2结合了无谄媚和无依赖与有刺激(一致性 = 0.915,原始覆盖度 = 0.479);总体解一致性达到0.886,覆盖度为0.567。两条路径虽有重叠但仍互为替代方案,且有一个要素同时出现在两者中:无谄媚。没有任何自主学习者组态能容忍谄媚的AI。低自主性遵循其自身逻辑,而非镜像另一侧。组态L1将缺失刺激与存在懒惰和依赖配对(一致性 = 0.905,原始覆盖度 = 0.539);L2将存在谄媚与缺失刺激和存在依赖配对(一致性 = 0.907,原始覆盖度 = 0.501);解一致性 = 0.893,覆盖度 = 0.578。依赖同时出现在低自主性组态中,缺失刺激亦然,而懒惰仅在其中一条路径中起作用,揭示出对称模型无法捕捉的因果不对称性(Ragin, 2008)。
表4
| 条件 | A1(高) | A2(高) | L1(低) | L2(低) |
|---|---|---|---|---|
| PAS | ⊗ | ⊗ | ● | |
| PAIS | ● | ⊗ | ⊗ | |
| MCL | ⊗ | ● | ||
| AIDEP | ⊗ | ⊗ | ● | ● |
| 一致性 | 0.899 | 0.915 | 0.905 | 0.907 |
| 原始覆盖度 | 0.535 | 0.479 | 0.539 | 0.501 |
| 唯一覆盖度 | 0.089 | 0.032 | 0.077 | 0.040 |
| 解一致性 | 0.886 | 0.893 | ||
| 解覆盖度 | 0.567 | 0.578 |
高与低学习者自主性的fsQCA组态。
●,存在;⊗,缺失;空白,无关。所有条件均为核心条件(无逻辑余项;解一致)。锚点(完全隶属/交叉点/完全不隶属):PAS 4.50/3.15/2.01;PAIS 4.66/3.50/2.01;MCL 4.50/3.00/1.50;AIDEP 4.07/2.67/1.50;ALS 4.67/3.59/2.56。
稳健性与敏感性分析
一系列检验考察了这些结果的稳健性。对每一个内生构念进行控制后,没有任何系数变动超过 0.002,而一个性别多样性虚拟变量也没有带来任何变化。针对自主性四个预测变量的高斯 copula 项——之所以适用,是因为所有回归变量都偏离了正态分布——均不显著(最小 p = 0.054,对应 AIDEP)(Park and Gupta, 2012),表明在常规水平上不存在内生性,同时仍与在局限性部分讨论的关于可能存在相互影响的谨慎态度相一致。一个反向中介模型,其中 AIDEP 被设定为先于 MCL,也产生了统计上显著的序列间接路径(PAS:β = −0.025;PAIS:β = 0.016),尽管这些路径的幅度大约仅为按理论提出的顺序所获得结果的一半。由于两种顺序都拟合横截面数据,本研究设计无法唯一确定从 MCL 到 AIDEP 的先后顺序;因此,所提出的顺序依赖于理论推理和先前的纵向研究先例,而非横截面模型拟合,序列结果最好被解释为与假设的时间顺序相一致,而非对其的验证。构成不变性在性别之间以及 STEM 与非 STEM 领域之间均成立(所有 c ≥ 0.998)(Henseler et al., 2016);经 Bonferroni 校正后仍存留的一处路径差异是,STEM 学生中刺激对懒惰的抑制作用更强(β = −0.456 对 −0.236,p = 0.003),而谄媚—依赖路径上一个名义上的性别差异(β = 0.261 对 0.043,p = 0.014)未能通过校正。Bootstrap 上界使每个 HTMT 均低于 0.85(最大值 = 0.664),平行分析保留了一个单一的 ALS 因子,并且在 PRI 阈值 0.65 至 0.75 以及频率阈值 3–10 的范围内,没有任何针对高自主性的 fsQCA 解包含谄媚或依赖的存在。相比之下,每一个低自主性解都包含依赖以及刺激的缺失。
讨论
本研究追问的是一位永远和颜悦色的导师会对被辅导者产生什么影响,并以一条关联链作答:感知到的谄媚伴随着自我调节的松弛,松弛的调节伴随着依赖,而依赖伴随着自主性的降低;相反,感知到的智力反驳则沿着完全相同的链条朝保护性方向运行。三种分析逻辑在这一解读上达成一致,且每一种都补充了其他逻辑无法提供的条款。
本研究的理论贡献涵盖三个方面。首先,它表明AI谄媚是一种可在心理层面测量的现象。先前研究已证实语言模型会奉承和顺从(Sharma et al., 2024),且接触会改变用户(Cheng et al., 2026);缺失的是将这种接触带入日常学习的 learner-side 变量。PAS量表填补了这一空白。该量表通过反转组织逢迎的方向构建而成(Kumar and Beyerlein, 1991),在首次使用中表现良好:内部一致、收敛有效,且在实证上与刺激、懒惰和依赖相区分。这在理论上很重要,因为卸载决策取决于对辅助工具的主观感知特征,而非其客观特征(Risko and Gilbert, 2016)。模型的基准谄媚分数可能永远不会进入学生的元认知考量;但他们对模型举止的印象可以,而现在这种印象的进入可以被测量。
A conceptual question follows from this measurement choice. Sycophancy and intellectual stimulation are, in the first instance, properties of an AI system and its underlying training, whereas this study operationalizes them as learner perceptions. Two considerations motivate this decision. Theoretically, cognitive offloading is governed by a learner’s subjective appraisal of an external aid rather than by its objective characteristics (Risko and Gilbert, 2016), so the perceived interaction style is the proximal variable through which any system-level behavior could plausibly shape self-regulation. Empirically, the same objective model behavior may be experienced differently across users depending on prior expectations, task, and prompting style, so a perceptual measure captures variance that a single system-level benchmark would obscure. This choice nonetheless carries a cost: perceived sycophancy need not correspond to logged model behavior, learners may misattribute agreeableness, and socially desirable responding or limited metacognitive insight may color their reports. The PAS and PAIS measures are therefore best understood as indices of experienced interaction style rather than as audits of model behavior, and linking them to conversational telemetry remains an important step for validating the constructs and closing the perception–behavior gap.
Second, the study delivers a mechanism. The hypothesized serial mediation held in full: sycophancy and stimulation ran through the laziness-dependence relay with opposite signs, and the relay carried substantial indirect effects on autonomy over and above the two direct paths. This extends cognitive offloading theory at both ends of its chain. Upstream, it locates a manipulable trigger of offloading in the tone of the tool rather than in fixed user traits or task demands. Downstream, it shows within one model how offloading shortcuts consolidate into dependence, a step earlier studies implied but did not test together (Wang et al., 2026; Zhai et al., 2024), and that the consolidated state, not the momentary shortcut, is most strongly tied to lost autonomy. The finding also gives quantitative form to the distinction between effects with technology and effects of technology on the learner once it is removed (Salomon et al., 1991): perceived sycophancy was associated with the former alongside an apparent cost to the latter. One asymmetry refines the account: we found no credible evidence of a direct stimulation-dependence association, only the route through engagement, suggesting that challenge protects by re-recruiting regulation rather than by breaking habit head-on, exactly where the offloading ratchet locates the leverage.
The third contribution is the answer that only method triangulation could deliver. The symmetric model showed near mirror-image total effects of the two styles; NCA showed that these are not interchangeable currencies, because minimum doses of pushback, restraint in flattery, engagement, and independence from the tool are each necessary for high autonomy, with stimulation the tightest bottleneck (Dul, 2016); fsQCA showed that no sufficient recipe for autonomous learners contains high sycophancy, while the recipes for failure are not mirror images of the recipes for success (Fiss, 2011; Ragin, 2008). Nice-to-have and must-have name different causal claims, and this study offers, to our knowledge, one of the first necessity tests of that claim in the AI-and-learning literature. Intellectual pushback passed it.
Two cautions temper this reading. First, although PLS-SEM, NCA, and fsQCA address distinct questions—average net effects, necessity in degree, and sufficient combinations—all three analyses are estimated on the same cross-sectional, self-report dataset. Their agreement therefore reflects complementary analytical perspectives on one body of evidence rather than independent replication across data sources, and their convergence should not be read as triangulated confirmation of a causal mechanism. Second, the necessary conditions identified by NCA hold in degree rather than in kind and are bounded by the present sample: they describe the levels of sycophancy, stimulation, laziness, and dependence compatible with high autonomy among these AI-using students, in this institutional and cultural setting, and with the systems they happened to use (Dul, 2016). Whether adequate intellectual stimulation or restraint in sycophancy is necessary for autonomous learning more generally is an empirical question for replication across other populations, institutions, and AI systems, and the present results should not be extended into universal requirements. Read within these limits, the configurational findings remain instructive: the absence of sycophancy recurs in every configuration sufficient for high autonomy, whereas AIDEP and the absence of intellectual stimulation recur in the configurations associated with low autonomy—an asymmetry between the routes to success and the routes to failure that variable-centered models cannot express (Ragin, 2008).
The findings turn a general worry into a specific design requirement. Calibrated disagreeableness belongs in the design brief, alongside default tutor personas that note errors, ask why, and delay automatic approval of first drafts. This aligns with evidence that cognitive forcing functions reduce over-reliance on AI advice (Buçinca et al., 2021) and that friction is a feature of learning systems, not a defect (Bjork et al., 2013; Zohar et al., 2026). Students who reach the upper fifth of autonomy experience stimulation at roughly half the observable range, so an interaction style optimized purely for user comfort was associated with autonomy scores toward the lower end of the observed range.
The commercial current runs the other way: most users prefer and trust agreeable systems (Cheng et al., 2026), and learner preference is an unreliable guide to learning value (Kirschner and van Merriënboer, 2013). Engagement metrics will reward the very style this study indicts, so design should target constructive and interactive engagement instead (Chi and Wylie, 2014).
Course designers and instructors hold complementary levers. AI-literacy instruction can teach students to recognize machine flattery and prompt against it, requesting counterarguments, error hunts, and Socratic questioning rather than validation, since users otherwise absorb AI stances with little scrutiny (Krügel et al., 2023). Course design should preserve unaided practice and proven effortful techniques (Dunlosky et al., 2013), keeping AI a sparring partner inside a learner-led paradigm, not an oracle above it (Ouyang and Jiao, 2021). Students cannot be assumed to notice the problem themselves, so the noticing must be taught.
Policymakers and institutional buyers can shift the market. Procurement standards can require configurable pushback modes, disclosure of agreeableness tuning, and pre-deployment sycophancy audits, operationalizing recent experimental calls (Cheng et al., 2026) within AI-in-education governance frameworks (Holmes and Tuomi, 2022). A null-adjacent finding sharpens the point: usage frequency retained only a modest negative link once perceptions were included in the model, so capping minutes of AI use is not a substitute for governing its manners.
The design is cross-sectional, so the serial ordering of laziness before dependence rests on theory and published precedent rather than observed time; and reciprocal loops, such as low autonomy inviting further offloading, remain plausible. Consistent with this, a reversed-mediator specification also produced significant, if smaller, serial indirect paths, so the reported ordering is best read as consistent with the hypothesized sequence rather than as evidence of temporal precedence. All focal constructs are self-reports; procedural separation, clean method-variance diagnostics, and the divergent calibration index reduce but cannot erase that concern (Podsakoff et al., 2003), and perceived sycophancy is not yet matched to logged model behavior. The sample is a student convenience sample from a single context, and the autonomy instrument was developed in Western higher education (Macaskill and Taylor, 2010), so cultural norms around politeness may shape both the perception of machine flattery and its consequences. The PAS scale, though it performed well on one deployment, needs measurement-invariance evidence across cultures and languages, temporal stability checks, and independent replication.
Experiments should manipulate model agreeableness in graded doses to trace the causal response curve that survey data can only approximate. Telemetry-paired designs can test whether conversational logs predict perceived sycophancy, closing the perception–behavior gap. Longitudinal panels can watch offloading consolidate into dependence, and intervention studies can test whether critique-eliciting prompting habits inoculate against this negative serial pathway. Meta-analytic evidence shows human–AI combinations often trail the best solo performer (Vaccaro et al., 2024), and hybrid designs that deliberately hand regulation back to the learner may buffer the dependence pathway identified here (Järvelä et al., 2023; Molenaar, 2022). The stronger protective stimulation path observed among STEM students represents one such boundary worth targeted study. Pairing graded performance and retention with self-reported autonomy would close the loop this model opens.
Conclusion
A tutor that always agrees is not a neutral convenience. Among 542 students, higher perceived sycophancy was associated with MCL, laziness with dependence, and dependence with diminished autonomy, the capacity education exists to build, while perceived intellectual pushback ran the same relay in reverse. Three methods converged on one sentence: there is no configuration of an autonomous learner in which the machine’s flattery stays high. The findings hand designers a dial to turn, educators a curriculum topic, policymakers a procurement clause, and learners a skill worth practicing: the discipline of being disagreed with.
Statements
Data availability statement
The raw data supporting the conclusions of this article will be made available by the authors, without undue reservation.
Ethics statement
This study was reviewed and approved by the Research Ethics Committee (REC) of the University of Ha’il with the approval number [H-2026-147], dated [21/05/2026]. Also, this study implemented all the procedures involving “human participants” in accordance with the Helsinki Declaration 1964 and its later amendments, and also with the ethical standards of the institutional research committee. The authors followed these steps: (1) Informed consent was obtained from participants, and (2) the purpose of the survey was clearly explained to the participants. They had a clear idea of how their data would be used and the extent of their involvement. This allowed the participants to agree and voluntarily participate and provide honest feedback.
Author contributions
GD: Conceptualization, Data curation, Formal analysis, Investigation, Methodology, Resources, Visualization, Writing – original draft. MC: Conceptualization, Data curation, Formal analysis, Funding acquisition, Methodology, Software, Writing – original draft. SA: Data curation, Funding acquisition, Investigation, Project administration, Visualization, Writing – review & editing. MM: Data curation, Funding acquisition, Investigation, Project administration, Resources, Software, Supervision, Visualization, Writing – review & editing. SN: Data curation, Formal analysis, Funding acquisition, Investigation, Resources, Software, Supervision, Validation, Visualization, Writing – review & editing.
Funding
The author(s) declared that financial support was received for this work and/or its publication This research was funded by the Scientific Research Deanship at the University of Ha’il—Saudi Arabia through project number RG-24094.
Conflict of interest
The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
Generative AI statement
The author(s) declared that Generative AI was used in the creation of this manuscript. Generative AI tools (a large language model) were used solely to support language editing and proofreading of the manuscript. They were not used to generate, collect, or analyze data, produce study content, or draw scientific conclusions. The authors reviewed and verified the entire text and take full responsibility for the content of the manuscript.
Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.
Publisher’s note
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.
Supplementary material
The Supplementary material for this article can be found online at: https://www.frontiersin.org/articles/10.3389/fpsyg.2026.1957447/full#supplementary-material
References
1
AndreassenC. S.TorsheimT.BrunborgG. S.PallesenS. (2012). Development of a Facebook addiction scale. Psychol. Rep.110, 501–517. doi: 10.2466/02.09.18.PR0.110.2.501-517,
2
BeauchampM. R.BarlingJ.Zhen LiMortonK. L.KeithS. E.ZumboB. D.et al. (2010). Development and psychometric properties of the transformational teaching questionnaire. J. Health Psychol.15, 1123–1134. doi: 10.1177/1359105310364175,
3
BjorkR. A.DunloskyJ.KornellN. (2013). Self-regulated learning: beliefs, techniques, and illusions. Annu. Rev. Psychol.64, 417–444. doi: 10.1146/annurev-psych-113011-143823,
4
BuçincaZ.MalayaB.GajosK. Z. (2021). To trust or to think: cognitive forcing functions can reduce overreliance on AI in AI-assisted decision-making. Proc. ACM Hum.-Comput. Interact.5, 1–21. doi: 10.1145/3449287
5
ChenY.WangM.YuanS.ZhaoY. (2025). Development and validation of the conversational AI dependence scale for Chinese college students. Front. Psychol.16:1621540. doi: 10.3389/fpsyg.2025.1621540,
6
ChengM.LeeC.KhadpeP.YuS.HanD.JurafskyD.et al. (2026). Sycophantic AI decreases prosocial intentions and promotes dependence. Science391:eaec8352. doi: 10.1126/science.aec8352,
7
ChiM. T. H.WylieR. (2014). The ICAP framework: linking cognitive engagement to active learning outcomes. Educ. Psychol.49, 219–243. doi: 10.1080/00461520.2014.965823
8
CohenJ. (1988). Statistical Power Analysis for the Behavioral Sciences. 2nd Edn New York: Lawrence Erlbaum Associates.
9
DarvishiA.KhosraviH.SadiqS.GaševićD.SiemensG. (2024). Impact of AI assistance on student agency. Comput. Educ.210:104967. doi: 10.1016/j.compedu.2023.104967,
10
DijkstraT. K.HenselerJ. (2015). Consistent partial least squares path modeling. MIS Q.39, 297–316. doi: 10.25300/MISQ/2015/39.2.02
11
DizonJ. I. W. T.MendozaN. B.GaševićD.GanoticeF. A.Jr. (2026). Assessing AI-driven metacognitive offloading: initial development and validation of the metacognitive laziness scale. ECNU Rev. Educ.9, 1–12. doi: 10.1177/20965311261450994
12
DulJ. (2016). Necessary condition analysis (NCA): logic and methodology of “necessary but not sufficient” causality. Organ. Res. Methods19, 10–52. doi: 10.1177/1094428115584005
13
DulJ.van der LaanE.KuikR. (2020). A statistical significance test for necessary condition analysis. Organ. Res. Methods23, 385–395. doi: 10.1177/1094428118795272
14
DunloskyJ.RawsonK. A.MarshE. J.NathanM. J.WillinghamD. T. (2013). Improving students’ learning with effective learning techniques: promising directions from cognitive and educational psychology. Psychol. Sci. Public Interest14, 4–58. doi: 10.1177/1529100612453266,
15
FanW.ChengL.WangY.ZhaoQ.LiY. (2026). In-class AI use and attitudes among university students: the different mediating roles of cognitive relief and cognitive offloading. Behav. Sci.16:1014. doi: 10.3390/bs16061014,
16
FanY.TangL.LeH.ShenK.TanS.ZhaoY.et al. (2025). Beware of metacognitive laziness: effects of generative artificial intelligence on learning motivation, processes, and performance. Br. J. Educ. Technol.56, 489–530. doi: 10.1111/bjet.13544
17
FissP. C. (2011). Building better causal theories: a fuzzy set approach to typologies in organization research. Acad. Manag. J.54, 393–420. doi: 10.5465/amj.2011.60263120
18
FlavellJ. H. (1979). Metacognition and cognitive monitoring: a new area of cognitive-developmental inquiry. Am. Psychol.34, 906–911. doi: 10.1037/0003-066x.34.10.906
19
FornellC.LarckerD. F. (1981). Evaluating structural equation models with unobservable variables and measurement error. J. Mark. Res.18, 39–50. doi: 10.1177/002224378101800104
20
GerlichM. (2025). AI tools in society: impacts on cognitive offloading and the future of critical thinking. Societies15, 1–28. doi: 10.3390/soc15010006,
21
GreckhamerT.FurnariS.FissP. C.AguileraR. V. (2018). Studying configurations with qualitative comparative analysis: best practices in strategy and organization research. Strategic Organ.16, 482–495. doi: 10.1177/1476127018786487
22
HairJ. F.HultG. T. M.RingleC. M.SarstedtM. (2022). A primer on Partial Least Squares Structural Equation Modeling (PLS-SEM). 3rd Edn. Thousand Oaks, CA: SAGE
23
HairJ. F.RisherJ. J.SarstedtM.RingleC. M. (2019). When to use and how to report the results of PLS-SEM. Eur. Bus. Rev.31, 2–24. doi: 10.1108/ebr-11-2018-0203
24
HenselerJ.RingleC. M.SarstedtM. (2015). A new criterion for assessing discriminant validity in variance-based structural equation modeling. J. Acad. Mark. Sci.43, 115–135. doi: 10.1007/s11747-014-0403-8
25
HenselerJ.RingleC. M.SarstedtM. (2016). Testing measurement invariance of composites using partial least squares. Int. Mark. Rev.33, 405–431. doi: 10.1108/imr-09-2014-0304
26
HolmesW.TuomiI. (2022). State of the art and practice in AI in education. Eur. J. Educ.57, 542–570. doi: 10.1111/ejed.12533
27
JärveläS.NguyenA.HadwinA. (2023). Human and artificial intelligence collaboration for socially shared regulation in learning. Br. J. Educ. Technol.54, 1057–1076. doi: 10.1111/bjet.13325
28
JiaW.PanL.NearyS. (2025). Effect of GenAI dependency on university students’ academic achievement: the mediating role of self-efficacy and moderating role of perceived teacher caring. Behav. Sci. (Basel)15:1348. doi: 10.3390/bs15101348,
29
KasneciE.SesslerK.KüchemannS.BannertM.DementievaD.FischerF.et al. (2023). ChatGPT for good? On opportunities and challenges of large language models for education. Learn. Individ. Differ.103:102274. doi: 10.1016/j.lindif.2023.102274,
30
KirschnerP. A.van MerriënboerJ. J. G. (2013). Do learners really know best? Urban legends in education. Educ. Psychol.48, 169–183. doi: 10.1080/00461520.2013.804395
31
KockN. (2015). Common method bias in PLS-SEM: a full collinearity assessment approach. Int. J. e-Collab.11, 1–10.
32
KockN.HadayaP. (2018). Minimum sample size estimation in PLS-SEM: the inverse square root and gamma-exponential methods. Inf. Syst. J.28, 227–261. doi: 10.1111/isj.12131
33
KrügelS.OstermaierA.UhlM. (2023). Chatgpt’s inconsistent moral advice influences users’ judgment. Sci. Rep.13:4569. doi: 10.1038/s41598-023-31341-0,
34
KrugerJ.DunningD. (1999). Unskilled and unaware of it: how difficulties in recognizing one’s own incompetence lead to inflated self-assessments. J. Pers. Soc. Psychol.77, 1121–1134. doi: 10.1037/0022-3514.77.6.1121,
35
KumarK.BeyerleinM. (1991). Construction and validation of an instrument for measuring ingratiatory behaviors in organizational settings. J. Appl. Psychol.76, 619–627. doi: 10.1037/0021-9010.76.5.619
36
LittleR. J. A. (1988). A test of missing completely at random for multivariate data with missing values. J. Am. Stat. Assoc.83, 1198–1202. doi: 10.1080/01621459.1988.10478722
37
MacaskillA.TaylorE. (2010). The development of a brief measure of learner autonomy in university students. Stud. High. Educ.35, 351–359. doi: 10.1080/03075070903502703
38
MolenaarI. (2022). Towards hybrid human-AI learning technologies. Eur. J. Educ.57, 632–645. doi: 10.1111/ejed.12527
39
MuneerS.DastgeerG.QureshiM. I.AlshammariA. S. (2026a). Students' attitudes toward AI teaching assistants in education: considering the role of characteristics and perceptions in Ha'il, Saudi Arabia. Acta Psychol.262:106096. doi: 10.1016/j.actpsy.2025.106096
40
MuneerS.SinghA.ChoudharyM. H.AnumS. (2026b). Emotional creepiness and technological proficiency - understanding AI/GPT tools in the digitalization era. Humanit. Soc. Sci. Commun.13:1228. doi: 10.1057/s41599-026-07532-1
41
OuyangF.JiaoP. (2021). Artificial intelligence in education: the three paradigms. Comput. Educ. Artif. Intell.2:Article 100020. doi: 10.1016/j.caeai.2021.100020
42
PappasI. O.WoodsideA. G. (2021). Fuzzy-set qualitative comparative analysis (fsQCA): guidelines for research practice in information systems and marketing. Int. J. Inf. Manag.58:102310. doi: 10.1016/j.ijinfomgt.2021.102310,
43
ParkS.GuptaS. (2012). Handling endogenous regressors by joint estimation using copulas. Mark. Sci.31, 567–586. doi: 10.1287/mksc.1120.0718,
44
PodsakoffP. M.MacKenzieS. B.LeeJ.-Y.PodsakoffN. P. (2003). Common method biases in behavioral research: a critical review of the literature and recommended remedies. J. Appl. Psychol.88, 879–903. doi: 10.1037/0021-9010.88.5.879,
45
RaginC. C. (2008). Redesigning Social Inquiry: Fuzzy Sets and Beyond. Chicago, Illinois (Chicago, IL), USA: University of Chicago Press.
46
RichterN. F.SchubringS.HauffS.RingleC. M.SarstedtM. (2020). When predictors of outcomes are necessary: guidelines for the combined use of PLS-SEM and NCA. Ind. Manag. Data Syst.120, 2243–2267. doi: 10.1108/imds-11-2019-0638
47
RingleC. M.WendeS.BeckerJ.-M. (2024). SmartPLS 4 [Computer Software]. SmartPLS GmbH. SmartPLS GmbH.
48
RiskoE. F.GilbertS. J. (2016). Cognitive offloading. Trends Cogn. Sci.20, 676–688. doi: 10.1016/j.tics.2016.07.002,
49
SalomonG.PerkinsD. N.GlobersonT. (1991). Partners in cognition: extending human intelligence with intelligent technologies. Educ. Res.20, 2–9. doi: 10.2307/1177234
50
SharmaM.TongM.KorbakT.DuvenaudD.AskellA.BowmanS. R.et al. (2024). Towards understanding sycophancy in language modelsProceedings of the Twelfth International Conference on Learning Representations (ICLR 2024), Vienna, Austria.
51
ShmueliG.SarstedtM.HairJ. F.CheahJ.-H.TingH.VaithilingamS.et al. (2019). Predictive model assessment in PLS-SEM: guidelines for using PLSpredict. Eur. J. Mark.53, 2322–2347. doi: 10.1108/ejm-02-2019-0189
52
SparrowB.LiuJ.WegnerD. M. (2011). Google effects on memory: cognitive consequences of having information at our fingertips. Science333, 776–778. doi: 10.1126/science.1207745,
53
StormB. C.StoneS. M. (2015). Saving-enhanced memory: the benefits of saving on the learning and remembering of new information. Psychol. Sci.26, 182–188. doi: 10.1177/0956797614559285,
54
TurelO.SerenkoA.GilesP. (2011). Integrating technology addiction and use: an empirical investigation of online auction users. MIS Q.35, 1043–1061. doi: 10.2307/41409972
55
VaccaroM.AlmaatouqA.MaloneT. (2024). When combinations of humans and AI are useful: a systematic review and meta-analysis. Nat. Hum. Behav.8, 2293–2303. doi: 10.1038/s41562-024-02024-1,
56
WangJ. (2026). Cognitive offloading through digital tools and its relationship with critical thinking, task persistence, and learning depth. Front. Psychol.17:1781101. doi: 10.3389/fpsyg.2026.1781101,
57
WangW.WuY.FangJ.YangC.WenL. (2026). When cognitive offloading becomes dependence: how AI dependence mediates the pathway from academic stress to burnout and anxiety. BMC Psychol. 14, 1–11.
58
WardA. F.DukeK.GneezyA.BosM. W. (2017). Brain drain: the mere presence of one’s own smartphone reduces available cognitive capacity. J. Assoc. Consum. Res.2, 140–154. doi: 10.1086/691462
59
YangZ.DengH.JiangN. (2025). The impact mechanism of artificial intelligence dependence on college students’ innovation capability: an empirical study from China. Front. Psychol.16:1732837. doi: 10.3389/fpsyg.2025.1732837,
60
ZhaiC.WibowoS.LiL. D. (2024). The effects of over-reliance on AI dialogue systems on students’ cognitive abilities: a systematic review. Smart Learn. Environ.11:28. doi: 10.1186/s40561-024-00316-7
61
ZhangS.ZhaoX.ZhouT.KimJ. H. (2024). Do you have AI dependency? The roles of academic self-efficacy, academic stress, and performance expectations on problematic AI usage behavior. Int. J. Educ. Technol. High. Educ.21:34. doi: 10.1186/s41239-024-00467-0
62
ZhaoX.LynchJ. G.Jr.ChenQ. (2010). Reconsidering baron and Kenny: myths and truths about mediation analysis. J. Consum. Res.37, 197–206. doi: 10.1086/651257
63
ZimmermanB. J. (2002). Becoming a self-regulated learner: an overview. Theory Pract.41, 64–70. doi: 10.1207/s15430421tip4102_2
64
ZoharE.BloomP.InzlichtM. (2026). Against frictionless AI. Commun. Psychol.4:39. doi: 10.1038/s44271-026-00402-1,
Keywords
AI dependence, AI sycophancy, generative artificial intelligence, learner autonomy, metacognitive laziness
Citation
Choudhary MH, Dastgeer G, Ahmed SA, Mohsin M and Naseem S (2026) How AI sycophancy shapes learner autonomy in digitalized learning: the mediating roles of metacognitive laziness and AI dependence. Front. Psychol. 17:1957447. doi: 10.3389/fpsyg.2026.1957447
Received
03 August 2026
Revised
30 August 2026
Accepted
04 September 2026
Published
30 September 2026
Volume
17 - 2026
Edited by
Daniel H. Robinson, The University of Texas at Arlington College of Education, United States
Updates
Copyright
© 2026 Choudhary, Dastgeer, Ahmed, Mohsin and Naseem.
This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.
*Correspondence: Ghulam Dastgeer, g.dastgeer@uoh.edu.sa
Disclaimer
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.
来源:Frontiers in Psychology · frontiersin.org
猜你喜欢
- 眼动实验比较生成式AI、传统搜索与混合检索对职校学生来源核查与迁移表现的影响Frontiers in Psychology · 1 小时前
- 数字健康干预对冠心病患者生活质量、焦虑与抑郁疗效的网络元分析Frontiers in Psychiatry · 1 天前
- 强化CBT治疗强迫症的随机对照试验元分析Frontiers in Psychiatry · 1 天前
- Epic Cosmos 230万人电子病历研究:孤独症谱系障碍人群自杀未遂风险分布Frontiers in Psychiatry · 1 天前
- 研究用眼动、EEG 与语义差异量表考察 AI 生成中国水墨画的观看反应Frontiers in Psychology · 1 天前