眼动实验比较生成式AI、传统搜索与混合检索对职校学生来源核查与迁移表现的影响
Responsible use of generative AI in Chinese vocational education: eye-tracking evidence on source verification and immediate near-transfer across three information-seeking conditions
一项针对96名中国职业院校学生的三组眼动实验比较了仅用生成式AI、传统网页搜索与混合检索三种信息检索方式。仅用生成式AI的参与者完成任务最快、感知认知负荷最低、任务信任最高,但对来源证据的关注更少、交叉核查尝试更少、修改答案更少;混合检索的来源证据注视比例(27.9%)、交叉核查尝试率(78.1%)和即时近迁移得分(86.99)最高。
眼动实验对比三种检索方式,显示混合检索在来源核查与即时近迁移上的表现差异。
译文尚不完整,完整内容请切换到原文。
摘要
引言:
生成式人工智能(GenAI)能够简化职业学习中的信息筛选与整合,但轻松获取流畅答案可能会降低学生核查信息来源的倾向。本研究基于双过程视角,比较了仅使用GenAI、传统网络搜索和混合搜索三种条件下,学生在来源验证、任务信任、感知认知负荷以及条件支持的即时近迁移表现方面的差异。
方法:
96名职业院校学生参加了一项三组间被试的眼动追踪实验。按学年、学科门类和先前AI使用经验进行分层后,他们被随机分配到三种条件(每组n = 32)。眼动数据、屏幕记录的行为日志、问卷和任务得分通过组间比较、相关分析和分层回归进行分析。
结果:
仅使用GenAI的参与者最快完成两项限时任务,报告的感知认知负荷最低、任务信任最高,但对来源证据的关注较少,交叉验证尝试较少,修改答案的频率也较低。传统网络搜索涉及更多的多来源浏览和证据比较,但耗时最多。混合搜索显示出最高的来源证据注视比例(27.9%)、交叉验证尝试率(78.1%)和即时近迁移得分(86.99),并且在近迁移表现上优于其他两种条件。在调整搜索条件和先前工具经验后,来源证据注视比例(β = 0.28,p = 0.006)和交叉验证尝试(β = 0.31,p = 0.002)仍与近迁移表现呈正相关,而任务信任则不然。
结论:
三种搜索模式在效率、验证和任务表现之间表现出不同的权衡。仅使用GenAI的搜索缩短了完成时间并降低了整体感知认知负荷,而混合搜索在指定条件下与更多的来源核查和更好的即时近迁移表现相关。这些发现表明,当学习者继续验证证据并保留最终判断的责任时,GenAI可以支持问题构建和信息组织。
1 引言
在数字化转型加速和人工智能快速发展的背景下,生成式人工智能(GenAI)正在重塑教育中知识的获取方式、学习支持的提供方式以及能力的培养方式(Kasneci et al., 2023;Yan et al., 2024)。与传统网络搜索相比,GenAI辅助搜索通过提供综合性的、面向任务的回答,可以减少与迭代搜索或查询重构、结果筛选、多页面浏览和信息整合相关的操作负担(Kaiser et al., 2025;Cai and Tian, 2025)。它能够以相对较低的成本生成结构化回答,并帮助学习者快速形成问题表征和初步解决方案(Kasneci et al., 2023)。与此同时,它可能压缩学习赖以发生的多源比较和证据验证过程,从而增加认知依赖、接受不准确信息以及批判性判断力减弱的风险(Lee et al., 2025)。联合国可持续发展目标4呼吁提供包容和公平的优质教育以及终身学习机会,以支持个人和社会的可持续发展(OECD, 2026)。从可持续职业教育的角度来看,来源验证和学习迁移是高质量数字素养的重要组成部分。这在涉及法规、安全、程序标准、时效性信息或其他高后果决策的职业任务中尤为重要。在此类任务中,来源的可靠性和范围直接影响专业判断的质量。而学习迁移则表明学生能否在新的职业情境中应用可靠信息和验证策略。这些能力共同支持可持续发展目标4所强调的有效学习成果、与职业相关的技能以及终身学习能力。
因此,AI工具的教育价值密切取决于学生如何信任它们以及如何验证其输出(Fu and Weng, 2024)。传统网络搜索促使学习者进入多个网页、比较来源并整合信息,但时间和认知成本更高。混合搜索可能通过允许学习者根据任务需要在AI生成、关键词扩展、网络搜索、来源验证和答案修订之间灵活切换,从而保留两种模式的优势。这种组合可能在操作成本与证据判断之间提供一种可行的平衡(Boetje et al., 2026;Yen et al., 2024)。因此,不同的搜索模式不仅仅是可替代的工具;它们可能引发不同的信息处理路径、信任形式和迁移结果。
然而,工具层面的互补性并不能保证学习者会自愿承担验证的成本。当学习者的领域知识有限、难以识别错误,或认为额外核查带来的收益不大时,他们可能会节省认知资源,倾向于接受流畅、即时可得的答案。关于技术采纳的研究同样表明,努力期望、社会影响和便利条件会塑造学生使用GenAI的意愿(BinJwair, 2025)。关于混合搜索的研究必须区分在指定验证要求下的有效性与学习者在自然情境中的自发采纳。
在真实的学习情境中,这一问题可能演变为过度依赖的风险。核心问题不仅仅是AI和搜索工具是否都可用,还包括学习者何时被提示进行验证、验证是否成为实际作答过程的一部分,以及工具设计是否能保持学习者对证据和最终判断的认知投入。
文献中仍存在若干空白。首先,大多数研究聚焦于普通高等教育或宽泛的学习情境,对职业教育的关注相对较少,而职业教育以实际应用和技能迁移为核心(Sun and Tian, 2026)。其次,技术接受度、学习结果和用户态度往往被分开考察,导致搜索模式、任务信任、来源验证和学习迁移之间的过程性联系未得到充分解释(García-Alonso et al., 2024;Chen et al., 2025)。在GenAI与传统搜索共存的学习环境中,以往研究也未能充分区分工具可供性与学习者实际验证行为。尽管混合搜索可能将生成效率与外部证据核查相结合,但关于学习者是否会自愿投入额外认知资源进行验证的直接证据仍然有限。实验要求下观察到的优势是否会转化为自然学习情境中的自发使用,也尚不清楚。
在此背景下,本研究以职业院校学生为对象,采用结合眼动追踪、行为日志、问卷和任务得分的多模态设计。在共同的、明确的验证要求下,参与者被随机分配到三种信息寻求条件之一。负责任的GenAI使用作为一个基于过程的解释框架,用于说明仅使用GenAI搜索、传统网络搜索和混合搜索如何影响来源相关注意、验证行为、整体感知认知负荷和即时近迁移表现。本研究探讨了三个研究问题。
RQ1:三种信息寻求条件在来源相关视觉注意和可观察的验证行为上有何差异?
RQ2:三种搜索条件在任务信任、整体感知认知负荷和条件支持的即时近迁移表现上有何差异?
RQ3:在控制搜索条件、先前AI使用经验和先前网络搜索经验后,对来源证据的注意和跨来源验证在多大程度上与即时近迁移表现相关?
通过回答这些问题,本研究将 GenAI 支持的学习研究从效率结果拓展到对信任、验证与迁移的基于过程的解释。这些发现也为 AI 在职业教育中的负责任整合提供了依据。这种整合应帮助学生从仅仅高效地使用 AI 转向负责任地使用 AI,同时支持高质量、可问责且可持续的职业教育。
2 文献综述
2.1 生成式 AI 的教育价值与风险及其带来的数字素养要求
关于 GenAI 支持的学习的研究已逐渐从证明工具有效性转向解释学习机制(Kasneci et al., 2023)。早期研究主要强调在问答、写作支持、反馈和资源组织方面的效率提升(Kasneci et al., 2023;Cai and Tian, 2025)。由于这些工具易于获取、交互性强且响应迅速,它们能够重塑学习支持,并为个性化学习、协作学习和即时辅导创造新机会。因此,GenAI 被界定为教育数字化转型的重要驱动力,并被纳入关于优质教育和可持续学习能力的讨论(OECD, 2026;Miao and Holmes, 2023)。较新的研究已将其教育价值从效率拓展到动机、任务理解和知识建构方面的变化(Yan et al., 2024)。AI 生成的内容可能提高可理解性、可及性和情境化支持,帮助学生组织问题、形成初步想法、检索文献、界定问题并识别可能的研究方向。
然而,GenAI 的教育价值不能仅凭其生成答案的速度来评判。AI 输出往往流畅、结构良好且即时。在缺乏充分人工监督的情况下生成的低质量 AI 内容——此处称为 AI 生成的“slop”——可能包含事实错误、缺失语境或来源不透明。其精致的外观可能使学习者更难区分可读性与可靠性(Kasneci et al., 2023;Lee et al., 2025),促使他们将表达质量等同于证据质量,并降低其验证动机(Anderl et al., 2024;Martín-Moncunill and Alonso Martínez, 2025)。关于 ChatGPT 等工具的研究既发现了教育机遇,也发现了风险,包括个性化反馈和学习支持,以及对学术诚信、批判性思维弱化和难以评估生成信息的担忧。因此,GenAI 支持的学习可能在提高效率的同时,也助长认知卸载、过度信任和来源忽视。
从数字素养的角度来看,关键问题不再仅仅是学生能否操作AI工具,而是他们能否在使用过程中保持判断、验证和迁移的能力。对于职业教育而言,数字素养应超越提示词构建和基本工具使用,涵盖来源敏感性、证据验证、偏见识别、元认知监控和负责任使用(Bråten et al., 2011;Walraven et al., 2009;Salmerón et al., 2020)。验证要求还应根据任务风险进行校准。涉及法规、工作场所安全、食品安全、消费者权益或财务判断的高后果任务,需要仔细核查来源出处、范围和时效性。低风险、开放式或创造性任务则不一定需要同等强度的验证(OECD, 2026;Miao and Holmes, 2023)。
现有研究已经确立了GenAI的学习支持价值和风险,但对其关注不足的是职业学生如何评估来源、校准信任,以及在真实任务中迁移所学内容。这一空白使信息搜寻环境本身成为问题的关键部分:不同的搜索支持条件可能会改变学习者在快速构建答案与来源验证之间的精力分配。因此,下一个问题是,仅使用GenAI、传统网络搜索和混合搜索是否会产生不同的信任、验证、认知需求和迁移模式。
2.2 不同搜索支持条件下的信任、信息验证与学习迁移
搜索支持条件会影响学生获取、评估和使用信息的方式。传统网络搜索要求学习者制定关键词、查看搜索结果、进入网页并比较来源。这一过程通常需要更多时间和认知努力,但也创造了接触多样信息、区分来源并做出更审慎判断的机会。相比之下,AI辅助检索以对话方式生成整合性回答,降低了结果筛选和信息组织的需求,使学习者能够更快地进入答案构建阶段。然而,同样的便利可能会压缩浏览、比较和验证过程,从而增加对现成答案的依赖。
任务信任为理解搜索模式如何影响学习提供了重要视角。适当程度的信任使学生能够接受工具支持并推进任务。传统搜索可能不会产生更强的主观信任,但它要求学习者通过点击来源、访问网页和跨页面比较来建立判断基础。混合搜索提供了一条中间路径。学生可以使用AI来澄清问题、生成关键词或形成初步答案,然后使用搜索引擎来确认事实、验证来源并修订回答。
混合搜索的可用性并不意味着学习者会自发采用验证策略。领域知识有限的学习者可能被生成答案的流畅性和即时性所吸引,停留在低成本的信息路径上,或者无法识别会触发额外核查的错误和不确定性。混合搜索的教育价值不仅取决于两种工具的可用性,还取决于任务要求、验证提示和评估标准是否能激活来源意识,并将额外的认知努力导向关键事实和证据。
从学习迁移的角度来看,搜索模式的效果不能仅通过完成速度来评估。工具支持的表现还必须与持久的能力发展区分开来:GenAI 产生的表现提升不一定构成学习或技能习得(Fauth and González-Martínez, 2021;Melumad and Yun, 2025;Yan et al., 2025)。迁移要求学生理解信息背后的依据和规则,并在新任务中重新应用这种理解。依赖现成的 AI 答案可能提高即时效率,但不会产生更深的理解。本研究考察学习者是否能在保持其分配的支持条件的同时,将来源验证、证据比较和答案修订策略重新应用于结构相似的任务。这种表现不等同于长期能力发展。
尽管先前研究已分别确立了 AI 辅助检索、传统搜索、任务信任和学习迁移的重要性,但三种检索模式之间的直接比较仍然有限。特别是,很少有研究在同一任务框架内比较仅 GenAI、传统网络和混合搜索。比较这些条件可以确定结果是否存在差异,但仅凭结果差异无法显示学习者在搜索过程中如何将注意力分配给来源证据。因此,需要过程层面的测量来确定信任和迁移的差异是否伴随着来源导向注意和验证的差异。
2.3 信息验证与学习过程研究中的眼动追踪
来源验证既是一种信息评估行为,也是一种注意力分配过程。仅靠问卷和最终分数无法确定学生是否关注了来源线索,或在答案、证据和作答区域之间反复移动。眼动追踪记录任务期间的视觉活动,包括注视持续时间、注视次数、AOI 停留比例、回视次数以及区域间的转换。这些指标可以揭示在信息搜索、证据判断和答案构建过程中注意力如何分布,从而提供与来源验证相关的过程层面证据。这些可供性使眼动追踪特别适用于研究数字环境中的来源评估(Gottschling et al., 2019;Liu and Cui, 2025;Hessels et al., 2025)。
教育技术、数字阅读和界面评估方面的研究表明,眼动追踪非常适合分析复杂数字环境中的认知加工过程。研究通常定义AOI,并考察注视如何在信息区域之间分布和转移,以刻画注意优先级、加工深度和任务策略。源区域的注视时间、回访和跨区域转移与信息验证尤为相关。数字素养较强的学习者更可能检查作者身份、来源、发布日期、引用线索和支持性证据的位置,而不是只关注主要内容。相反,注意力集中在答案或内容区域,而来源线索大多被忽视,可能表明来源意识有限、证据判断较浅。这一解释沿袭了先前关于在线来源评估和基于注视的可信度评估指标的研究(Gottschling and Kammerer, 2021;Guan and Lin, 2025;Tsai et al., 2022)。
当GenAI进入学习环境时,信息寻求不再是线性的网页浏览序列。AI答案生成、搜索结果检查、来源评估和答案修订可能并行发生,或反复循环。眼动追踪可以揭示学习者是否主要停留在AI回答区域、是否将注意力重新导向来源证据,或是否在生成的答案、网页内容、搜索结果和回答区域之间反复移动。当与屏幕录制、行为日志、问卷和任务得分结合时,它可以支持对搜索条件下视觉注意、外显行为、主观体验和学习结果进行整合性解释。多模态学习分析可以通过将注视与记录的动作和表现证据进行三角互证,增强这一推断(Järvelä et al., 2021;Mohammadi et al., 2025)。
尽管眼动追踪研究已经揭示了数字阅读和信息评估中的注意分配,但在GenAI支持的学习中,多模态过程证据仍然有限。整合视觉注意指标、行为日志、主观测量和迁移任务得分,可以揭示注意力分配在哪里,以及验证行为是否发生。然而,这些测量本身并不能解释为什么流畅的AI输出可能促使人快速接受,或者为什么验证提示可能触发更审慎的检查。因此,解释此类过程差异需要一种关于启发式加工和分析式加工如何被调用的解释性说明。
2.4 理论框架:双过程视角下的AI依赖与验证
为提供这一解释性说明,本研究以双过程、启发式—分析式视角作为主要框架。双过程理论认为,快速、自动的类型1加工通常会生成默认反应,而依赖工作记忆的类型2加工则在需要更审慎判断时更充分地参与(Evans and Stanovich, 2013)。在人工智能辅助决策中,Buçinca et al. (2021) 进一步表明,用户可能依赖对人工智能的泛化启发式信任,而非分析性地评估每一条建议。认知强制功能可以减少这种过度依赖,尽管它们也会增加感知到的努力。
从这个视角来看,生成式人工智能输出的速度、流畅性和整合呈现可以减少操作摩擦,同时增加启发式信任并减少自发验证(Lee and See, 2004;Vasconcelos et al., 2023)。相反,外部证据、验证提示以及混合搜索的可供性可能会在关键主张处引入有限、有目的的“认知摩擦”,促使学习者比较来源、修改答案并进行更多分析性加工。目标不是最大化认知负荷,而是将认知努力从机械性搜索操作转向必要的证据判断。信任校准和负责任的人机交互作为互补概念,用于解释学习者是否会根据证据调整信任并保留对最终判断的责任。它们不被视为额外的独立理论体系。
在此框架内,信息寻求条件是实验操纵。对来源—证据AOI的关注、交叉验证尝试和答案修改是关键过程指标。任务信任和整体感知认知负荷表征主观状态,而条件支持的即时近迁移表现是主要学习结果。
3 方法论
本研究采用三条件被试间实验设计,并整合了来自眼动追踪、屏幕录制、行为日志、问卷和任务得分的多模态数据。在数据清洗和信度检查之后,使用描述性统计、组间检验、相关分析和回归分析来考察不同信息寻求条件下的来源相关注意、验证行为、任务信任、整体感知认知负荷和即时近迁移表现。负责任的生成式人工智能使用作为基于过程的解释框架,而非测量的复合变量。它通过任务信任并结合对来源的关注、证据验证、答案修改和条件支持的即时近迁移表现来解释。研究路径如图1所示。
图1
3.1 研究设计与参与者
3.1.1 研究设计、招募与样本量
本研究采用了一项三条件被试间对照实验,参与者为中国某职业学院的在读学生。参与者于2026年5月9日至6月9日期间招募,构成便利样本。他们的年龄为18–25岁,来自机电工程与制造、财经商贸、物流管理及相关专业等职业领域。纳入标准包括具备基本的中文阅读能力、基本的网络搜索技能、视力正常或矫正后正常,以及能够成功完成眼动仪校准。
目标样本量依据眼动追踪研究的实际限制条件以及有效眼动数据的预期保留率确定。它也反映了在三种实验条件之间进行比较所需的最低参与者人数(Lakens, 2022)。由于眼动追踪实验对实验室环境、校准质量、每位参与者的测试时间以及有效注视数据的比例有相对严格的要求,本研究设计为每种条件至少保留30名有效参与者。额外招募参与者是为了应对校准失败、退出、方案违规以及有效注视数据率低于75%的情况。共招募了102名学生。在六名被排除的参与者中,四人的有效注视数据低于预先设定的75%阈值,一人因严重近视或其他眼部相关困难导致无法可靠校准而未能完成眼动仪校准,一人因在未获允许时使用AI而违反了所分配的信息搜寻方案。没有参与者因退出而被排除。每名被排除的参与者均按主要排除原因计一次。最终样本包含96名参与者,GenAI-only条件、传统网络搜索条件和混合搜索条件各分配32名。
进行了敏感性功效分析,以评估最终样本是否充分支持主要组间比较。对于涉及三组、α = 0.05、统计功效为0.80的单因素方差分析,总样本N = 96足以检测约Cohen’s f = 0.322的效应,对应于η2 = 0.094。因此,该样本提供了足够的功效来检测中等或更大程度的组间效应,尽管其识别小效应的能力仍然有限(Kang, 2021)。
3.1.2 分层与随机分配
为了提高不同条件下在教育背景和先前使用 AI 工具经验方面的可比性,参与者按学习年限、学科类别和先前 AI 使用经验进行了分层。学习年限分为一年级、二年级或三年级。学科类别包括机电工程与制造、金融与商务、物流管理及相关领域。先前 AI 使用经验根据实验前问卷的得分分为低、中、高三个等级。纳入来自不同年级和学科背景的学生,拓宽了对职业教育中学习阶段和领域的覆盖范围。采用分层随机化,使这些特征在不同条件之间尽可能均匀分布,并减少其对搜索行为的潜在混杂影响。低、中、高 AI 经验类别同样仅作为随机化前的分层变量,不构成额外的处理条件。在每个层内,参与者通过计算机生成的置换区组随机化被分配到三个实验条件之一。随机化由一名未参与任务评分的研究助理执行,结果评估者在评分期间对分组分配保持盲态。置换区组随机化遵循了用于平衡处理组的既定分配原则(Lim and In, 2019)。
3.1.3 伦理与知情同意
在正式实验开始前,已获得辽宁经济职业技术学院马克思主义学院的伦理批准。批准编号为 LECVTC-MY-2026-001,批准日期为 2026 年 5 月 8 日。本研究按照《赫尔辛基宣言》中关于涉及人类参与者研究的伦理原则进行。在参与之前,所有学生均获得了关于研究目的、实验程序、所收集数据类型、数据匿名化以及退出权利的信息。所有参与者均签署了书面知情同意书。
3.2 实验工具与检索条件
为提高可比性,所有实验均在同一个实验室进行,使用相同的显示器、屏幕分辨率、浏览器设置和网络配置。任务材料、时间限制、作答要求和验证目标在三种条件下完全相同。所有参与者均被要求评估信息、核查相关主张,并在必要时修改答案。各条件仅在可用的信息获取工具以及由此产生的获取来源可供性方面有所不同。
3.2.1 标准化环境与仅使用生成式 AI 条件
在仅使用生成式人工智能(GenAI)的条件下,参与者使用 ChatGPT Plus Instant,界面语言设置为中文,并选择 GPT-5.3 作为标准化模型。整个实验过程中使用同一个测试账号,每次会话前均禁用记忆功能或清除所有历史对话。不允许访问传统搜索引擎和外部网页。参与者可以进行多轮对话,请求解释或示例,提出追问,并指示系统澄清或重新组织其回答。他们不得使用传统搜索引擎或独立访问外部网页。因此,在该条件下,来源核查仅限于 GenAI 界面内显示的引用标识、来源卡片、超链接、域名标签或归属信息(如有)。
3.2.2 传统网络搜索条件
传统网络搜索条件下的参与者使用 Microsoft Edge 148.0 版本和 Bing 搜索引擎。实验在中国大陆进行,使用本地 Bing 搜索服务。在数据收集期间,实验室所使用的大陆版 Bing 未提供 Copilot Search/生成式回答功能。我们还审查了传统网络搜索条件的屏幕录像,确认 Bing 搜索结果页面上未显示任何 AI 生成的答案面板、生成式摘要或 Copilot Search 输出。因此,参与者与标准 Bing 搜索结果进行交互,并独立访问外部网页,未接触到 Bing 的 AI 生成答案层。每次会话前,清除浏览器缓存、搜索历史和登录状态,以尽量减少个性化推荐对搜索结果的影响。参与者可以独立制定关键词查询、浏览搜索结果、打开网页、比较来源并综合出答案,但不得使用任何生成式人工智能工具。
3.2.3 混合搜索条件与跨条件可比性
混合搜索条件下的参与者可以同时使用上述 GenAI 工具和搜索引擎。GenAI 界面和 Bing 在单独的浏览器标签页中打开,参与者根据其信息需求手动切换标签页;两个界面并非并排显示。因此,界面切换和回访应在手动切换浏览器标签页的背景下理解。他们可以根据需要使用任一工具来构建问题框架、生成关键词或初步回答、核实来源、确认事实性主张并修改答案。任务材料、时间限制、回答要求和实验室环境在三种条件下均保持一致;条件之间的差异仅在于参与者可获得的信息搜寻支持形式。
3.3 任务开发与实验流程
3.3.1 任务开发、专家评审与试点测试
在预备阶段,依据风险敏感原则开发了三个与职业情境相关的候选任务。这些任务涉及电商售后规则、餐饮服务中的食品安全以及设备电池安全;每个任务都涉及规则适用性、时效性或安全后果等问题,因而需要验证。为减少因任务材料类型差异可能带来的混淆,四位职业教育与教育技术领域的专家对候选任务进行了评审,并对15名学生进行了试点测试。筛选标准包括任务难度、职业相关性、语言复杂度、可验证性和完成时间。最终选择电商售后政策任务用于正式实验。所有三种条件均在相同时间限制下完成同一任务。
3.3.2 前测与眼动校准
每次实验环节持续约40至60分钟。参与者到达实验室后,首先阅读并签署知情同意书。随后,他们完成一份实验前问卷,内容涵盖人口统计信息、先前AI使用经验、网络搜索经验、对AI的接受度以及基线信任水平。之后进行五点眼动仪校准。
3.3.3 信息整合任务
在主要任务中,参与者在各自被分配的信息寻求条件下完成一项职业情境化的信息整合任务,时间限制为20分钟。任务指定了一个专业角色、一个现实的工作相关场景、相互冲突或不完整的信息以及相关的操作约束。参与者需要查找和评估信息、验证相关主张,并提出适当的修改或解决方案。任务全程记录眼动和屏幕活动。
3.3.4 即时近迁移任务
紧接着,参与者在相同的指定信息寻求条件下完成一项10分钟的即时近迁移任务。即时近迁移被定义为将来源验证、证据比较和答案修订策略应用于一个新的但结构相似的职业问题。迁移任务改变了案例主题、任务目标、冲突证据模式和决策约束,同时保留相同的职业领域和相当的难度水平。因此,它旨在评估即时近迁移,而非延迟迁移、跨领域远迁移或无技术支持的独立迁移(Forsyth, 2018)。
在迁移任务中保持指定条件不变,使得能够考察验证策略如何在不同信息寻求环境中被重新应用。因此,所得分数应被解释为条件支持下的即时近迁移表现,而非无搜索或GenAI支持的独立表现。相应地,该测量不被视为长期“能力发展”的证据,而是被操作化为验证相关表现以及在结构相似任务中验证策略的即时应用。
3.3.5 后测与数据导出
完成两项任务后,参与者填写了一份实验后问卷。随后,研究人员导出了眼动追踪数据,整理了屏幕录制内容,并收集了参与者的书面回答。实验流程和多模态数据采集工作流见图2。
图2
3.4 眼动追踪设备、AOI定义与多模态测量
3.4.1 眼动追踪设备与实验室设置
眼动追踪数据使用Tobii Pro Spark屏幕式眼动仪采集,采样率为60 Hz。使用Tobii Pro Lab 1.232进行记录回放、数据预处理和兴趣区(AOI)标注。每位参与者在正式实验前均完成了眼动仪校准。所有实验均在稳定光照条件下进行,屏幕分辨率和显示缩放比例保持恒定(Carter and Luke, 2020;Orquin and Holmqvist, 2018)。
3.4.2 功能性AOI框架
由于GenAI交互页面、搜索引擎结果页面和外部网页在视觉结构上存在显著差异,AOI的界定依据功能对等性而非固定的空间坐标。研究区分了四种页面类型:任务页面、GenAI交互页面、搜索结果页面和外部来源网页。任务页面包含任务说明区和回答输入区。GenAI交互页面包括提示词输入区、GenAI生成回答区和来源或引用提示区。传统网页搜索界面包括搜索结果区、网页内容区和来源信息区。
为进行跨条件分析,功能上可比的区域被聚合为四个更高层级的类别:任务理解AOI、信息获取AOI、来源证据AOI和答案构建AOI。这种分类使得在页面布局和界面结构存在差异的情况下,仍能对等效信息功能的注意力进行比较(Hooge et al., 2026)。
3.4.3 动态页面状态与AOI归一化
这些界面可滚动,并在交互过程中动态变化(Blascheck et al., 2017)。因此,AOI根据每个记录的页面状态中可见的内容进行编码,而非被视为永久固定的屏幕坐标。每当导航、滚动或新生成的内容改变可见布局时,就标记一个新的页面状态。AOI边界相应更新,注视时长在所有归属于同一功能性AOI类别的页面状态中累计。可见视口之外的内容被排除在注视时间计算之外。这一程序使得对功能等效信息区域的注意力能够在页面切换、滚动事件和界面布局变化中进行聚合。
根据每个界面中显示的信息来定义来源线索。在 GenAI 界面上,这些线索包括可见的引文标记、来源卡片、超链接、发布者或域名标签,以及关于生成性主张来源的明确陈述。在搜索结果页面上,来源线索包括结果标题、URL 或域名、发布日期、发布者或机构标识,以及结果摘要中呈现的来源相关信息。在外部网页上,来源线索包括作者、机构或发布者信息、发布日期、URL 或域名、参考文献列表,以及与官方政策、法规或文件相关的标识符。
当给定页面状态中未显示任何来源信息时,该状态不编码来源证据 AOI。因此,来源证据 AOI 的缺失被视为一种界面特征,而非参与者忽略了可用来源信息的证据。AOI 定义在主要分析之前即已确立,并使用相同的功能性编码协议一致地应用于所有记录。
为减少访问页面数量和总浏览时间差异的影响,AOI 指标根据每位参与者的总有效注视时长或有效观看时间(视情况而定)进行归一化处理。来源证据 AOI 内注视时长的比例,通过将该 AOI 内的注视时长除以参与者的总有效注视时长来计算。跨页面的注视数据按功能等效的 AOI 类别进行汇总,而非基于单个页面布局或固定屏幕位置进行比较。
3.4.4 眼动追踪指标与解释边界
主要的眼动追踪指标包括总注视时长、注视次数、各 AOI 内的停留比例、回视次数以及 AOI 之间的转换。这些指标刻画了整体注意力分配、对来源证据的注意、重复查看以及功能区域之间的切换。然而,这些指标本身仅表明视觉注意和观看行为;它们无法确定参与者是否理解了证据或进行了批判性评估。因此,这些指标与行为日志、交叉验证、答案修改和任务表现结合进行解释。AOI 定义的示意图见 图 3。
图 3
3.4.5 行为日志编码与验证定义
行为日志来自连续屏幕录制和人工编码。编码变量包括提示或查询的数量、页面切换、来源点击、访问的外部网页、是否发生交叉验证、两个限时任务的总完成时间、直接复制 AI 输出,以及最终回答中引用的来源数量。总完成时间指标为 20 分钟信息整合任务与 10 分钟即时近迁移任务的活跃完成时间之和,最大可能值为 30 分钟。两名编码者使用标准化方案独立对录制内容进行编码;分歧通过讨论解决,并报告了编码者间信度。为避免在“验证”这一笼统标签下混淆不同层次的行为,我们区分了来源访问、验证尝试和独立跨来源验证。来源点击和网页访问主要表示来源访问。由于仅使用生成式 AI 条件下的参与者无法独立访问外部网页,要求生成式 AI 为同一事实主张提供第二个来源被编码为跨来源验证尝试。在传统网络搜索和混合搜索条件下,当参与者实际访问并比较至少两个相互独立的外部来源时,被编码为独立跨来源验证。因此,在三种条件下报告的比例描述的是广义、特定可供性意义上的跨来源验证尝试;其实现方式并不等同于独立外部验证。
3.4.6 问卷测量与学习结果
实验前和实验后问卷使用五点李克特式量表收集人口统计学信息,并测量 AI 接受度、任务信任、总体感知认知负荷、感知验证努力和元认知监控。主要分析前使用 Cronbach's α 评估内部一致性(Taber, 2018)。学习结果评分涵盖信息准确性、论证完整性、来源质量和迁移应用质量。认知负荷在两个任务后的任务后问卷中测量一次,并操作化为总体感知认知负荷,反映在指定搜索与回答工作流程中体验到的一般心理努力和任务处理负担。得分越高表示总体感知负荷越大。该测量未区分内在负荷、外在负荷或与学习相关的认知投入。
学习成果包括信息整合任务和即时近迁移任务的表现。回答从信息准确性、论证完整性、来源质量和迁移应用质量四个方面进行评价。每个维度采用五点量表评分,各维度得分合并计算总分。两名评分者使用标准化评分量规独立评价所有回答,并在正式评分前完成校准培训。评分结束后计算评分者间信度。当信度达到预设标准时,取两次评分的均值作为最终得分。存在重大分歧的案例由第三名评分者复核。研究变量、相应指标、数据来源及其在分析中的作用汇总于表1。
表1
| 变量类别 | 核心指标 | 数据来源 | 分析作用 |
|---|---|---|---|
| 眼动追踪测量 | 总注视时长、注视次数、来源-证据AOI停留比例、回视次数和AOI转换次数 | 从Tobii Pro Lab导出的数据 | 刻画注意力分配、对来源信息的注意以及功能区域之间的切换 |
| 行为日志 | 提示/查询、页面切换、来源点击、外部网页访问、交叉验证尝试、复制、答案修改以及两项限时任务的总完成时间 | 屏幕录制和人工编码 | 刻画搜索路径、验证深度和对技术工具的依赖 |
| 问卷测量 | 人口统计学特征、AI接受度、任务信任、认知负荷、感知验证努力和元认知监控 | 任务前后问卷 | 评估主观判断、感知认知需求和验证倾向 |
| 学习成果 | 信息整合任务质量和即时近迁移表现 | 任务回答和评分量规 | 评估学习成果和验证策略的即时应用 |
变量、指标、数据来源和分析作用。
3.5 数据准备与统计分析
3.5.1 数据预处理与质量控制
对预定义兴趣区(AOI)导出眼动追踪数据并进行质量筛查。如果校准失败、观察到明显的注视漂移,或有效注视数据比例低于预设阈值75%,则排除参与者。在招募样本中,排除情况包括有效注视数据不足75%(n = 4)、因严重近视或其他眼部困难导致眼动仪校准失败(n = 1),以及涉及未经授权使用AI的方案违规(n = 1)。行为日志由两名编码者独立编码,任务回答由两名评分者独立评价。分别计算编码者间信度和评分者间信度。随后按参与者ID合并眼动追踪测量、行为指标、问卷回答和任务表现得分,创建最终分析数据集。
3.5.2 总体分析策略与假设检验
统计分析使用 IBM SPSS Statistics 26.0 版进行。对主要变量计算描述性统计量,并在进行推断性检验前检查分布特征和方差齐性。满足参数分析假设的变量采用单因素方差分析(ANOVA),随后进行适当的事后比较。当这些假设不满足时,使用 Kruskal-Wallis H 检验。分类变量,包括交叉验证的发生情况,采用卡方检验进行分析,当期望单元格频数不足时则采用 Fisher 精确检验。对于每一项推断性检验,以 α = 0.05 作为决策阈值。p < 0.05 的值被视为足以拒绝相应原假设的统计证据,而 p ≥ 0.05 则被视为不足以拒绝原假设。结果酌情以频数和百分比或 M ± SD 的形式报告,并在适用时附上检验统计量、p 值、效应量和置信区间。SPSS 26.0 支持所需的假设检验、ANOVA、分类检验和非参数检验、相关性分析、分层回归以及诊断分析,并作为主要的统计软件包。热图和扫描路径仅用于支持对视觉搜索策略的描述性解释;它们未被用作显著性检验的主要依据。
3.5.3 RQ1:视觉注意与验证行为
针对 RQ1,分析考察了不同条件之间在来源相关视觉注意和可观察验证行为上的差异。主要的视觉注意指标为来源-证据 AOI 注视比例、对来源相关区域的回访,以及功能不同 AOI 之间的转换。行为指标包括请求支持性证据、引用使用、复制行为、答案修改和交叉验证。由于不同条件下获取外部信息的途径不同,来源点击和外部网页访问作为条件特定的描述性指标报告。三条件比较首先刻画了信息寻求工作流程之间的总体行为差异。在仅 GenAI 条件下,“跨来源验证尝试”指就同一主张向 GenAI 请求第二个来源,而在两个启用网络的条件中,它可能涉及独立的外部跨来源验证。后者在仅限于传统网络搜索和混合搜索条件的补充比较中单独考察。
3.5.4 RQ2:主观测量与任务表现
针对RQ2,在任务信任、整体感知认知负荷和即时近迁移表现上对三种条件进行了比较。根据情况使用单因素方差分析或Kruskal Wallis H检验,在总体检验显著后进行事后比较。为避免将效率与验证质量或学习结果混为一谈,操作效率仅由两项限时任务中的合并主动完成时间和整体感知认知负荷来定义。验证则分别由来源-证据AOI停留时间、回访、跨来源验证尝试和答案修订来代表;任务表现由信息整合和即时近迁移得分来代表。任务信任作为单独的主观判断变量报告,未纳入操作效率的定义。感知验证努力、元认知监控和信息整合表现作为次要结果报告,用以补充RQ2中提到的三个构念。
3.5.5 RQ3:相关、分层回归与敏感性分析
针对RQ3,使用零阶相关和分层回归来检验来源相关注意和交叉验证是否与即时近迁移表现相关。搜索条件、先前AI使用经验和先前网络搜索经验在焦点过程指标之前进入模型,从而在控制条件和两种信息寻求工具的先前熟悉度后,估计它们与近迁移表现的关联。使用容忍度和方差膨胀因子(VIF)评估预测变量之间的多重共线性,以确认来源点击、交叉验证和其他预测变量可以同时纳入最终模型。
由于各条件对独立外部来源的访问并不等同,涉及交叉验证的回归结果结合各条件的可供性进行解释。当使用独立外部跨来源验证作为预测变量时,在两个可联网条件内进行了敏感性分析。回归系数被解释为统计关联,而非因果效应的证据。此外,任务信任和整体感知认知负荷均在两项任务后的同一任务后评估中测量,而验证行为则来自过程记录。由于该设计未建立因果中介所需的时间顺序,我们未估计以任务信任为中介变量的因果中介模型。相反,使用零阶相关和多变量回归来描述整体感知认知负荷、任务信任、验证行为和即时近迁移表现之间的统计关联。
4 结果
4.1 数据质量、测量信度与基线等效性
4.1.1 数据质量与测量信度
最终分析纳入96名参与者,其中GenAI-only搜索、传统网络搜索和混合搜索条件各分配32名。在保留的参与者中,有效注视数据的平均比例为88.64%,SD = 5.73%,超过了预设的75%阈值。主要问卷量表的Cronbach's α系数范围为0.81至0.88,而行为编码的Cohen's κ值范围为0.82至0.91。信息整合和即时近迁移任务评分的评分者间信度分别为ICC = 0.89和ICC = 0.87。主要连续变量中未发现严重离群值,其分布特征和方差齐性总体上适合后续分析。
4.1.2 各条件间的基线等效性
基线平衡检验显示,三种条件在性别、年龄、既往AI使用经验、网络搜索经验、AI接受度或基线任务信任方面均无显著差异,所有p > 0.05,支持分层随机化后的广泛可比性。为提高透明度,表2还呈现了全样本及各条件下学习年份、学科集群和低/中/高既往AI使用分层的分布;这些分层不构成额外的实验组。尽管如此,既往AI使用经验和网络搜索经验仍作为协变量保留在回归模型中,以考虑个体对这两种信息寻求工具熟悉程度的差异。
表2
| 变量 | 总计(N = 96) | GenAI-only(n = 32) | 传统网络搜索(n = 32) | 混合搜索(n = 32) | 检验统计量 | p |
|---|---|---|---|---|---|---|
| 总计 | 96 | 32 | 32 | 32 | ||
| 性别,男/女 | 45/51 | 15/17 | 14/18 | 16/16 | χ2 = 0.25 | 0.883 |
| 年龄,岁 | 20.39 ± 1.20 | 20.31 ± 1.18 | 20.47 ± 1.26 | 20.38 ± 1.21 | F = 0.14 | 0.872 |
| 学习年份,n | χ2 = 0.13 | 0.998 | ||||
| 第一年 | 33 | 11 | 11 | 11 | ||
| 第二年 | 32 | 11 | 10 | 11 | ||
| 第三年 | 31 | 10 | 11 | 10 | ||
| 学科集群,n | χ2 = 0.13 | 0.998 | ||||
| 机电 | 33 | 11 | 11 | 11 | ||
| 金融/商务 | 32 | 10 | 11 | 11 | ||
| Logistics/related fields | 31 | 11 | 10 | 10 | ||
| Prior AI-use experience | 3.24 ± 0.76 | 3.24 ± 0.76 | 3.18 ± 0.81 | 3.31 ± 0.74 | F = 0.22 | 0.806 |
| Prior AI-use category, n | χ2 = 0.19 | 0.996 | ||||
| Low | 31 | 10 | 10 | 11 | ||
| Moderate | 34 | 12 | 11 | 11 | ||
| High | 31 | 10 | 11 | 10 | ||
| Web-search experience | 3.77 ± 0.68 | 3.72 ± 0.68 | 3.81 ± 0.71 | 3.78 ± 0.66 | F = 0.13 | 0.879 |
| AI acceptance, pretest | 3.45 ± 0.62 | 3.46 ± 0.62 | 3.39 ± 0.65 | 3.51 ± 0.59 | F = 0.31 | 0.734 |
| Baseline task trust | 3.27 ± 0.57 | 3.28 ± 0.57 | 3.22 ± 0.61 | 3.30 ± 0.55 | F = 0.17 | 0.844 |
Participant characteristics and baseline equivalence across conditions.
Continuous variables are reported as M ± SD.
4.2 Source related visual attention and verification behavior
4.2.1 Source-related visual attention
To address RQ1, source-related visual attention and observable verification behavior were compared across the three information-seeking conditions. Total fixation duration did not differ significantly, whereas fixation count, source-evidence AOI dwell proportion, source-related revisits, and cross-AOI transitions did. The main statistical results are presented in Table 3.
Table 3
| Measure | GenAI only | Conventional web search | Hybrid search | Test statistic | p | Effect size |
|---|---|---|---|---|---|---|
| Panel A. Eye tracking measures | ||||||
| Total fixation duration (s) | 1,098 ± 118a | 1,123 ± 129a | 1,110 ± 121a | F = 0.33 | 0.718 | ηp2 = 0.007 |
| Fixation count | 2,850 ± 420b | 3,220 ± 460a | 3,380 ± 450a | F = 12.01 | <0.001 | ηp2 = 0.205 |
| Source evidence AOI dwell proportion (%) | 11.8 ± 7.1c | 22.7 ± 8.0b | 27.9 ± 8.6a | F = 34.41 | <0.001 | ηp2 = 0.425 |
| Source related revisit count | 3.2 ± 2.1c | 7.1 ± 3.0b | 9.4 ± 3.3a | F = 38.81 | <0.001 | ηp2 = 0.455 |
| AOI transition count | 18.6 ± 8.2c | 42.4 ± 13.1b | 50.7 ± 14.6a | F = 58.96 | <0.001 | ηp2 = 0.559 |
| Panel B. Behavioral measures | ||||||
| Total completion time across the two timed tasks (min) | 24.2 ± 3.4c | 29.1 ± 4.1a | 26.8 ± 3.7b | F = 13.72 | <0.001 | ηp2 = 0.228 |
| Source clicks | 0.9 ± 0.8b | 4.8 ± 1.5a | 5.7 ± 1.8a | F = 101.95 | <0.001 | ηp2 = 0.687 |
| External webpages visited | Not available | 5.5 ± 2.2a | 4.9 ± 2.1a | t = 1.12 | 0.269 | d = 0.28 |
| Cross-verification attempt, n (%) | 8 (25.0)b | 22 (68.8)a | 25 (78.1)a | χ2 = 21.03 | <0.001 | V = 0.468 |
| Direct copying, n (%) | 14 (43.8)a | 0 (0.0)b | 4 (12.5)b | χ2 = 21.33 | <0.001 | V = 0.471 |
| Answer revision, n (%) | 10 (31.3)b | 19 (59.4)ab | 25 (78.1)a | χ2 = 14.48 | 0.001 | V = 0.388 |
| Sources cited in final response | 1.3 ± 1.0b | 3.4 ± 1.4a | 4.1 ± 1.5a | F = 39.12 | <0.001 | ηp2 = 0.457 |
Eye tracking and behavioral measures across the three information access conditions.
Continuous variables are shown as M ± SD and categorical variables as n (%). Within each row, values that do not share a letter differ at the multiplicity adjusted p < 0.05 level. Source clicks and external webpage visits are condition dependent affordance indicators, the external webpage comparison was restricted to the two web enabled conditions.
The clearest pattern concerned attention to source evidence: source-evidence dwell and related revisits and transitions were lowest in the GenAI-only condition and highest in the hybrid-search condition. GenAI-only participants concentrated more attention within the information-acquisition region, whereas hybrid-search participants shifted more often between content, source evidence, and response construction. Figure 4 provides the corresponding AOI distribution.
Figure 4
4.2.2 Observable verification behavior
Behavioral logs showed the same overall contrast (Table 3). GenAI-only participants had the shortest combined active completion time across the two timed tasks but showed less cross-verification and answer revision and more direct copying, whereas hybrid search maintained relatively high verification and revision with less time than conventional web search. Because external-source access differed by condition, source clicks and webpage visits are best interpreted as affordance-dependent measures rather than pure indicators of verification motivation. In the GenAI-only condition, the 25.0% cross-verification-attempt rate specifically denotes requests to GenAI for a second source for the same claim; independent external comparison was possible only in the web-enabled conditions.
The exploratory temporal plot was consistent with this pattern: GenAI-only participants moved quickly to the AI interface, conventional web-search participants progressed toward external webpages, and hybrid-search participants alternated more often among the AI interface, search results, and source webpages (Figure 5).
Figure 5
4.3 Questionnaire measures and task performance
4.3.1 Overall group differences and pairwise comparisons
To address RQ2, the three conditions were compared in terms of task trust, overall perceived cognitive load, perceived verification effort, metacognitive monitoring, information-integration performance, and immediate near-transfer performance. Omnibus tests were significant for all six outcomes; complete descriptive statistics, effect sizes, and adjusted pairwise comparisons are reported in Table 4.
Table 4
| Variable | GenAI only | Conventional web search | Hybrid search | F (2, 93) | p | ηp2 |
|---|---|---|---|---|---|---|
| M ± SD | ||||||
| Task trust | 4.18 ± 0.49a | 3.42 ± 0.58c | 3.74 ± 0.53b | 16.3 | <0.001 | 0.26 |
| Cognitive load | 3.21 ± 0.57b | 3.91 ± 0.55a | 3.48 ± 0.52b | 13.33 | <0.001 | 0.223 |
| Perceived verification effort | 3.12 ± 0.62b | 3.76 ± 0.54a | 4.01 ± 0.50a | 21.84 | <0.001 | 0.32 |
| Metacognitive monitoring | 3.35 ± 0.56c | 3.71 ± 0.52b | 4.02 ± 0.47a | 13.41 | <0.001 | 0.224 |
| Information integration task score | 78.34 ± 7.16b | 80.72 ± 7.88ab | 84.41 ± 6.92a | 5.57 | 0.005 | 0.107 |
| Immediate near transfer task score | 76.99 ± 8.25b | 81.65 ± 8.20b | 86.99 ± 7.45a | 12.61 | <0.001 | 0.213 |
Questionnaire measures and task performance across the three information access conditions.
Questionnaire variables were measured on 5-point Likert type scales, task scores were transformed to a 100 point scale. Means within a row that do not share a letter differ significantly in Tukey HSD comparisons at p < 0.05.
The key contrasts were as follows. GenAI-only search produced the highest task trust and the lowest overall perceived cognitive load, whereas hybrid search produced the highest metacognitive monitoring and the strongest immediate near-transfer performance. Perceived verification effort was highest under hybrid search but did not differ significantly from conventional web search. Hybrid search significantly outperformed GenAI-only search on both information-integration and near-transfer performance and also outperformed conventional web search on near-transfer performance. Overall perceived cognitive load did not differ significantly between the GenAI-only and hybrid conditions.
4.3.2 The high-trust-low-verification pattern
The reliability-relevant process indicators revealed a high-trust–low-verification pattern in the GenAI-only condition. Although task trust was highest, source-evidence attention, cross-verification, and answer revision were lower and direct copying was higher than in the hybrid-search condition (Tables 3 and 4). In hybrid search, trust was lower but verification and revision were substantially more frequent. The 25.0% GenAI-only cross-verification-attempt rate denotes requests for a second AI-provided source rather than independent external-source verification. Taken together, high subjective trust did not ensure substantive evidence processing.
4.4 Correlations and hierarchical regression
4.4.1 Zero-order correlations
Zero-order correlations showed that immediate near-transfer performance was positively related to source-evidence AOI dwell proportion (r = 0.46, p < 0.001) and cross-verification attempts (r = 0.52, p < 0.001), and negatively related to overall perceived cognitive load (r = −0.31, p = 0.002). Task trust was not significantly related to near-transfer performance (r = −0.14, p = 0.172), but it was negatively related to cross-verification attempts (r = −0.38, p < 0.001), as reported in Table 5. Because source-access measures were shaped by condition-specific affordances, these correlations are descriptive; the hierarchical regression provides the primary adjusted estimates. The analyses concern observable verification behavior rather than a separately measured self-report construct of need for source verification.
Table 5
| Variable | M | SD | 1 | 2 | 3 | 4 | 5 | 6 | 7 |
|---|---|---|---|---|---|---|---|---|---|
| 1. Trust | 3.78 | 0.61 | 1 | ||||||
| 2. Source dwell | 20.8 | 10.9 | −0.29** | 1 | |||||
| 3. Source clicks | 3.8 | 2.5 | −0.34** | 0.42** | 1 | ||||
| 4. Cross-verification attempt | 0.573 | 0.497 | −0.38** | 0.48** | 0.56** | 1 | |||
| 5. Cognitive load | 3.53 | 0.61 | −0.22* | 0.18 | 0.25* | 0.12 | 1 | ||
| 6. Metacognition | 3.69 | 0.58 | −0.30** | 0.41** | 0.36** | 0.45** | −0.26* | 1 | |
| 7. Near transfer | 81.88 | 8.9 | −0.14 | 0.46** | 0.39** | 0.52** | −0.31** | 0.44** | 1 |
Descriptive statistics and zero-order correlations among process measures and immediate near-transfer performance.
*p < 0.05; **p < 0.01; Trust, task trust; Source dwell, source evidence AOI dwell proportion; Metacognition, metacognitive monitoring; Near transfer, immediate near transfer score. Cross verification was coded 0/1, its correlations with continuous variables are point biserial correlations.
4.4.2 Hierarchical regression
Table 6 presents the hierarchical regression results. The final model explained 51% of the variance in immediate near-transfer performance (adjusted R2 = 0.46). After adjustment for search condition and prior tool experience, hybrid search (β = 0.24, 95% CI [0.04, 0.44], p = 0.019), source-evidence AOI dwell proportion (β = 0.28, 95% CI [0.08, 0.48], p = 0.006), and cross-verification attempts (β = 0.31, 95% CI [0.12, 0.50], p = 0.002) were positively associated with near-transfer performance, whereas overall perceived cognitive load was negatively associated with performance (β = −0.20, 95% CI [−0.38, −0.02], p = 0.031). Source clicks were not independently associated with performance once the more substantive verification indicator was included.
Table 6
| Predictor | Model 1β | Model 2β | Model 3β | Approx. 95% CI for β | p |
|---|---|---|---|---|---|
| Prior AI-use experience | 0.13 | 0.08 | 0.04 | [−0.12, 0.20] | 0.624 |
| Web-search experience | 0.19 | 0.11 | 0.06 | [−0.10, 0.22] | 0.462 |
| Conventional web search | 0.21* | 0.09 | [−0.10, 0.28] | 0.352 | |
| Hybrid search | 0.43*** | 0.24* | [0.04, 0.44] | 0.019 | |
| Task trust | −0.12 | [−0.30, 0.06] | 0.178 | ||
| Source-evidence AOI dwell proportion | 0.28** | [0.08, 0.48] | 0.006 | ||
| Source clicks | 0.13 | [−0.04, 0.30] | 0.141 | ||
| Cross-verification attempt | 0.31** | [0.12, 0.50] | 0.002 | ||
| Overall perceived cognitive load | −0.20* | [−0.38, −0.02] | 0.031 | ||
| R2 | 0.07 | 0.26 | 0.51 | ||
| Adjusted R2 | 0.05 | 0.23 | 0.46 | ||
| ΔR2 | 0.07 | 0.19*** | 0.25*** | ||
| F | 3.51* | 7.99*** | 9.95*** |
Hierarchical regression predicting immediate near-transfer performance.
The GenAI-only condition was the reference category. *p < 0.05; **p < 0.01; ***p < 0.001.
4.4.3 Web-enabled sensitivity analysis
In the web-enabled sensitivity analysis (n = 64), independent cross-source verification remained positively associated with immediate near-transfer performance (β = 0.29, 95% CI [0.05, 0.53], p = 0.021), whereas source clicks were not significant (β = 0.12, 95% CI [−0.12, 0.36], p = 0.318). This result reduces, but does not eliminate, the interpretive problem created by unequal source-access affordances across all three conditions.
4.5 Descriptive eye tracking visualizations
To illustrate the visual search patterns under the three information-seeking conditions, one participant was selected from each condition whose primary eye tracking measures were close to the central tendency of the corresponding group. Heatmaps and scanpath visualizations were then generated for these selected cases. Participant selection was based jointly on the proportion of dwell time within the source-evidence AOI, the number of revisits to source-related regions, and the number of transitions between AOIs, thereby reducing the influence of atypical cases on the visual interpretation.
4.5.1 Heatmap patterns across conditions
The representative heatmaps showed distinct patterns of visual attention on the task page across the three information-seeking conditions. In the conventional web-search condition, fixations were concentrated mainly in the upper portion of the task-material area, whereas attention to the task-requirement area was relatively limited and dispersed. In the GenAI-only condition, fixations were more strongly concentrated within the task-requirement area, indicating a more localized pattern of attention on the task page. By contrast, the hybrid-search condition showed broader visual coverage across the task-material area while retaining a distinct concentration of fixations within the task-requirement area. Overall, the representative heatmaps suggest more localized task-page attention in the GenAI-only and conventional web-search conditions and broader visual coverage in the hybrid-search condition (Figure 6).
Figure 6
4.5.2 Scanpath patterns across conditions
The scanpath further indicated that gaze in the GenAI-only condition remained more often within a single content region, with relatively few transitions among functional regions. The conventional web-search condition showed more switching between search results and webpage content, whereas the hybrid-search condition showed more frequent cross-region revisits, particularly among the generated response, source information, and webpage content. These descriptive visualizations were consistent in direction with the group-level eye-tracking and behavioral results, but they were used only to illustrate possible information-processing pathways, as shown in Figure 7.
Figure 7
5 Discussion
The findings indicate that the three information-seeking conditions involved trade-offs among operational efficiency, evidence processing, trust, and task performance rather than a single mode that was superior on every outcome. GenAI-only search was the fastest and produced the lowest overall perceived cognitive load. Conventional web search was the most demanding, but it preserved more extensive multi-source browsing. Hybrid search had intermediate operational costs and was associated with greater source attention, more answer revision, and better immediate near-transfer performance. Task trust likewise did not align uniformly with performance. It was highest in the GenAI-only condition, which had the lowest near-transfer score, lowest in the conventional web-search condition, and intermediate in the hybrid-search condition. Responsible use in vocational education should be evaluated jointly in terms of operational cost, evidence processing, trust calibration, and task performance.
5.1 The efficiency–verification trade-off
GenAI reduced the operational effort associated with iterative searching or query reformulation, result screening, multi-page navigation, and information organization, but may also have compressed source tracing, evidence comparison, and judgment of applicability. Eye-tracking and behavioral results showed less source-evidence AOI dwell, fewer source revisits and cross-region transitions, and lower rates of cross-verification and answer revision in the GenAI-only condition. Thus, lower overall perceived cognitive load cannot be equated with higher-quality learning. For combined active completion time across the two timed tasks, hybrid search required only 2.6 min more than GenAI-only search (26.8 vs. 24.2 min), and the difference in overall perceived cognitive load was not significant (3.48 vs. 3.21). Yet source-evidence dwell proportion increased from 11.8 to 27.9%, cross-verification attempts from 25.0 to 78.1%, and answer revision from 31.3 to 78.1%, while direct copying fell from 43.8 to 12.5%. Relative to conventional web search, hybrid search saved 2.3 min in combined active completion time across the two timed tasks, produced lower overall perceived cognitive load (3.48 vs. 3.91), and produced a higher source-evidence dwell proportion (27.9% vs. 22.7%). Its advantage was not maximal efficiency, but a more favorable verification–performance trade-off at moderate operational cost.
From a dual-process perspective, these differences are better understood as differences in how effort was allocated than as a simple contrast between high and low cognitive load. The rapid and coherent output of GenAI may encourage heuristic acceptance with relatively little cognitive effort, whereas source access, explicit verification requirements, and comparison across tools can introduce limited analytic processing at critical evidence checkpoints. Buçinca et al. (2021) similarly found that cognitive forcing functions reduced uncritical reliance on AI advice while increasing perceived effort. The educational objective should not be to maximize cognitive load, but to direct limited cognitive resources toward source comparison, evidential judgment, and answer revision.
This interpretation is consistent with Stadler et al. (2024), who reported that LLMs can reduce the burden of information seeking even when lower cognitive effort is accompanied by shallower reasoning. The higher load associated with conventional web search may include both operational inefficiency and productive effort devoted to evidence evaluation. Hybrid search instead enables a division of labor in which AI supports problem framing and preliminary organization while web search is used to verify key facts and conditions of applicability. Verification can thus be preserved without restoring all the operational costs of conventional search. Rather than eliminating friction, educational design should retain limited, purposeful cognitive friction at critical evidence checkpoints so that AI output remains provisional information subject to verification.
The study measured overall perceived cognitive load and did not distinguish intrinsic load, extraneous load, or learning-relevant cognitive investment. These results do not establish that hybrid search incurred no additional cognitive cost, nor can overall perceived load be treated as a direct indicator of deep processing. Recent review evidence similarly suggests that the effects of GenAI on cognitive load depend on scaffolding, intensity of use, prior knowledge, and task design (Qian et al., 2026). In the present sample, overall perceived cognitive load was not linearly associated with attention to source evidence or cross-source verification attempts. A more plausible interpretation is that some cognitive resources were reallocated from search operations to source verification, evidence comparison, and answer revision. Moreover, the formal task explicitly required participants in all conditions to locate, evaluate, and verify information, while source-access channels differed by condition. Accordingly, the higher verification rate in the hybrid-search condition reflects the combined influence of the verification requirement and tool affordances and should not be equated with a spontaneous verification tendency in naturalistic use.
5.2 Trust calibration, verification, and transfer
Task trust did not produce a linear benefit in this study. Trust was highest in the GenAI-only condition (4.18 ± 0.49), whereas immediate near-transfer performance was lowest (76.99 ± 8.25). In the hybrid-search condition, trust was lower (3.74 ± 0.53) and near-transfer performance was highest (86.99 ± 7.45). At the individual level, task trust was not significantly correlated with near-transfer performance (r = −0.14, p = 0.172). By contrast, source-evidence AOI dwell proportion and the affordance-specific cross-verification-attempt variable remained positively associated with near-transfer after adjustment for condition and prior experience. Reliable use depends not on maximizing trust, but on adjusting trust to the sufficiency and results of verification. The risk of AI-generated “slop” lies not in its AI-generated status alone, but in the mismatch between trust and evidential quality when low-quality content is presented fluently and without adequate sourcing.
Lee et al. (2025) found that greater confidence in GenAI among knowledge workers was associated with less critical thinking and that, in GenAI-supported work, critical thinking shifted toward information verification, response integration, and task management. The current results extend that finding through eye tracking, behavioral logs, and task scores. Source-evidence AOI dwell proportion, revisits, and transitions captured attention to evidence and movement among functional regions, while cross-verification and answer revision indicated whether judgments changed in response to evidence. These measures, however, cannot independently establish critical evaluation or deep understanding.
The process measures also distinguished source access from source verification. Source clicks no longer explained unique variance after more substantive indicators were included, whereas attention to source evidence and cross-verification attempts remained associated with near-transfer performance. Opening a webpage does not establish that its evidence was understood or used. For this reason, we distinguished source access, cross-source verification attempts, and independent external cross-source verification. In the GenAI-only condition, the 25.0% verification rate represented requests for a second AI-provided source concerning the same claim; independent external verification was possible only in the web-enabled conditions. The value of hybrid search lay in enabling learners to move through a sequence of generating, tracing, comparing, and revising information and to repeat that sequence in a structurally similar task. Its advantage lay primarily in the verification process and immediate performance, not operational efficiency or overall trust. AI literacy should extend beyond tool operation to source sensitivity, evidential judgment, metacognitive monitoring, and dynamic trust calibration (Long and Magerko, 2020; Lee and See, 2004).
5.3 Theoretical, design, and educational implications
Theoretically, the findings identify an important boundary condition for a dual-process account of AI reliance. Lower overall perceived cognitive load does not automatically indicate a better learning process, just as greater load does not necessarily indicate deeper processing. What matters is whether uncertainty or high-consequence claims trigger analytic engagement with evidence. The absence of a stable positive association between task trust and immediate near-transfer performance likewise suggests that responsible human-AI collaboration should prioritize trust calibrated to the available evidence rather than simply increasing trust in AI.
These findings frame responsible use as a process that can be observed and shaped through design: AI can support generation and organization, while learners retain responsibility for verification, revision, and final judgment. This division of labor is consistent with human-centered AI principles emphasizing human control, system reliability, and retained accountability (Amershi et al., 2019; Shneiderman, 2020). In vocational education, verification should be risk sensitive. Regulatory, safety-related, procedural, time-sensitive, and other high-consequence information should enter a more stringent verification workflow, whereas low-risk, open-ended, or creative tasks may be supported through lighter prompts or selective checks (OECD, 2026; Miao and Holmes, 2023).
Systems can organize information around explicit claim–evidence relationships. Traceable sources can be placed near factual, policy-related, and quantitative claims and can identify the issuing organization, publication date, location in the original material, and scope of the evidence. AI responses and web-based evidence can be presented side by side, and the system can record source inspection, evidential judgment, and answer revision to create a traceable evidence chain. Source cues alone, however, do not demonstrate that verification has occurred. If citations serve primarily as signals of credibility, learners may infer that cited content is necessarily reliable. Hybrid-search designs should combine uncertainty cues, verification of key claims, comparison of independent sources, and revision in response to new evidence, concentrating additional effort at high-risk or high-uncertainty points.
In instruction, GenAI output can be framed as a provisional proposal rather than a model answer. Tasks involving conflicting sources, changing rules, contextual constraints, or evidential gaps can require students to identify provenance, specify conditions of applicability, and justify whether a recommendation should be accepted. A pre-submission verification prompt can ask students to flag unverified claims and document how their answer changed in response to evidence. Bastani et al. (2025) showed that unconstrained GenAI support could improve performance while the tool was available yet undermine subsequent independent performance, whereas learning safeguards mitigated this effect. Assessment should likewise extend beyond the final product to the evidence-processing pathway. Criteria can include source quality, claim–evidence alignment, cross-source checking, recognition of conflicts, justification of revisions, and communication of uncertainty. In this way, GenAI use can move from maximizing efficiency toward a balanced allocation of tool support and learner responsibility for judgment.
5.4 Recommendations and implementation priorities
At the immediate course and assessment level, instructors should treat GenAI output as provisional rather than authoritative. For high-consequence factual, regulatory, numerical, or procedural claims, assignments can require a brief claim–evidence log, verification against an authoritative and current source, and a note explaining how the evidence changed the final answer. Verification prompts should be placed at consequential evidence checkpoints rather than applied indiscriminately to every sentence. Grading criteria should reward source quality, claim–evidence alignment, conflict recognition, justified revision, and communication of uncertainty.
At the medium-term curriculum and system level, vocational institutions should embed risk-sensitive source verification, trust calibration, and documentation of AI-supported decisions within digital-literacy modules and program assessments. Curriculum leaders and teacher-development teams can prepare discipline-specific verification cases, while system and interface designers can provide traceable provenance, publication dates, evidence scope, uncertainty cues, and side-by-side comparison of generated claims with primary sources. Institutions should evaluate these practices through repeated tasks, delayed assessments, and tool-withdrawal checks before treating short-term assisted performance as durable capability development.
5.5 Limitations and future research
Several limitations define the scope of the findings. First, the sample came from a single vocational college (N = 96). Although the experimental task approximated a vocational context, it simplified extended learning and authentic workplace demands and included an explicit verification requirement that may have elevated verification behavior above everyday GenAI use. Because the task was rule oriented and carried plausible real-world consequences, the value of verification observed here should not be generalized directly to low-risk, open-ended, or creative tasks. Second, external-source affordances differed across the three conditions, meaning that differences in cross-verification reflect both behavioral choice and tool availability. The sensitivity analysis within the two web-enabled conditions reduced this confounding but could not decompose the complete workflow effect into pure tool and motivational components. The study also did not directly manipulate AI-generated “slop,” model hallucination rates, or source traceability. It therefore provides evidence about verification behavior and trust calibration, not a direct test of learners’ ability to identify low-quality generated content.
Third, the immediate near-transfer task retained the assigned support condition and therefore measured tool-supported transfer rather than independent performance after tool withdrawal. Without a tool-withdrawal test, delayed follow-up, or repeated measurement, the findings cannot support claims about long-term capability development. Eye tracking primarily measures visual attention, while screen recordings and logs capture observable behavior; neither can independently establish comprehension, deep processing, or critical evaluation. Interpretation therefore relied on converging evidence from gaze, cross-verification, answer revision, and task scores. Task trust and overall perceived cognitive load were also measured at the same post-task assessment, precluding the temporal ordering required for mediation analysis. Accordingly, the study did not test a causal pathway from cognitive load through task trust to source verification. Finally, behavioral intention, effort expectancy, facilitating conditions, and other adoption factors were not measured. The study therefore evaluates execution under assigned conditions rather than the mechanisms governing spontaneous adoption in natural use (BinJwair, 2025). The study also did not measure self-efficacy, need for cognition, epistemic vigilance, or habitual dependence on a particular information-seeking mode; these unmeasured individual differences may influence both tool choice and verification behavior.
Future research should include a wider range of institutions, disciplines, tasks, and AI tools and incorporate delayed and tool-withdrawal assessments. It should also measure overall perceived cognitive load, trust, and subsequent verification dynamically over time. Experimental work should also manipulate verification-prompt intensity, task risk, generated-content quality, and source traceability to examine spontaneous adoption of hybrid strategies, retention of verification behavior, and trust calibration. Combining eye tracking with interviews and classroom-process data may further clarify the mechanisms involved. Larger prospectively designed studies should additionally model self-efficacy, need for cognition, epistemic vigilance, and habitual tool dependence as potential moderators of verification and transfer. Where mediation is theoretically proposed, future studies should establish temporal ordering through repeated or experimentally separated measurements before estimating indirect effects.
6 Conclusion
This study compared three information-seeking workflows. GenAI-only search, conventional web search, and hybrid search. The results revealed trade-offs among operational efficiency, evidence processing, and task performance. GenAI-only search was the fastest and produced the lowest overall perceived cognitive load, but it was accompanied by less attention to source evidence, fewer cross-source verification attempts, and fewer answer revisions. Conventional web search required the most time and produced the greatest perceived load, while preserving more multi-source browsing and evidence comparison. Hybrid search had intermediate operational costs and produced greater source-related attention, more answer revision, and better condition-supported immediate near-transfer performance. It is more appropriately described as producing a favorable verification–performance profile under the present experimental conditions, rather than as the most efficient method. Because source-access affordances differed across conditions, the observed differences reflect the combined effects of complete workflows and tool availability and cannot be attributed solely to verification motivation.
Task trust showed no stable positive relationship with immediate near-transfer performance. By contrast, the proportion of dwell time within the source-evidence AOI and the affordance-specific cross-verification-attempt variable remained significantly associated with near-transfer after controlling for search condition and prior tool experience. Source clicks showed no independent association. The same pattern appeared in the sensitivity analysis restricted to the two web-enabled conditions. These results sharpen the distinction between source access and source evaluation: opening a source is not equivalent to verifying it. What appears more transferable is comparing evidence concerning the same factual claim across sources and revising one’s judgment accordingly. From a dual-process perspective, the more informative mechanism is whether analytic processing is triggered at critical evidence checkpoints, not whether overall trust or overall perceived cognitive load is simply maximized or minimized.
For vocational education, these findings support a conditional model of responsible GenAI use. When verification requirements are activated and tool affordances permit, AI can support problem framing and information organization while learners retain responsibility for source checking, evidential judgment, and final decisions. Instructional and system design should reduce barriers to initiating verification through traceable sources, prompts to check critical claims, and aligned assessment requirements. Verification intensity should be calibrated to task risk rather than imposed uniformly across all tasks.
The conclusions are bounded by the single-institution sample, short-duration tasks, explicit verification requirement, and unequal source-access affordances. Because the immediate near-transfer task retained the assigned tools, the findings should not be generalized to delayed transfer, far transfer, unsupported performance, or long-term capability development. Future research should test the relationships among verification strategies, trust calibration, and learning transfer across institutions, vocational domains, and interface designs, including delayed and tool-withdrawal transfer tasks.
Statements
Data availability statement
The original contributions presented in the study are included in the article/supplementary material, further inquiries can be directed to the corresponding author.
Ethics statement
The studies involving humans were approved by the Institutional Review Board (School of Marxism, Liaoning Economy Vocational and Technical College, Shenyang, China). The studies were conducted in accordance with the local legislation and institutional requirements. The participants provided their written informed consent to participate in this study.
Author contributions
JM: Validation, Conceptualization, Writing – review & editing, Supervision, Methodology, Funding acquisition, Formal analysis, Project administration. HY: Investigation, Writing – review & editing, Resources, Visualization. XL: Visualization, Resources, Validation, Data curation, Writing – original draft, Conceptualization, Software, Methodology.
Funding
The author(s) declared that financial support was received for this work and/or its publication. This work was supported by the Liaoning Provincial Social Science Planning Fund (grant no. L21DSZ008).
Acknowledgments
The authors would like to thank the participating students and the research assistants who contributed to participant recruitment, experimental implementation, behavioral coding, and task-response evaluation.
Conflict of interest
The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
Generative AI statement
The author(s) declared that Generative AI was used in the creation of this manuscript. ChatGPT, developed by OpenAI, was used solely for English-language polishing, grammar checking, and readability improvement. Generative AI was not used to collect or generate the study data, conduct the statistical analyses, or determine the scientific findings and conclusions. All AI-assisted text was critically reviewed and revised by the authors, who take full responsibility for the accuracy, integrity, and final content of the manuscript.
Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.
Publisher’s note
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.
References
1
AmershiS.WeldD.VorvoreanuM.FourneyA.NushiB.CollissonP.et al (2019) Guidelines for human–AI interactionProceedings of the 2019 CHI Conference on Human Factors in Computing SystemsNew York, NY: Association for Computing Machinery
2
AnderlC.KleinS. H.SarigülB.SchneiderF. M.HanJ.FiedlerP. L.et al. (2024). Conversational presentation mode increases credibility judgements during information search with ChatGPT. Sci. Rep.14:17127. doi: 10.1038/s41598-024-67829-6,
3
BastaniH.BastaniO.SunguA.GeH.KabakcıÖ.MarimanR. (2025). Generative AI without guardrails can harm learning: evidence from high school mathematics. Proc. Natl. Acad. Sci. USA122:e2422633122. doi: 10.1073/pnas.2422633122,
4
BinJwairA. (2025). Predicting STEM students’ adoption of generative AI in academic contexts: an application of the UTAUT model. Front. Educ.10:1669750. doi: 10.3389/feduc.2025.1669750
5
BlascheckT.KurzhalsK.RaschkeM.BurchM.WeiskopfD.ErtlT. (2017). Visualization of eye tracking data: a taxonomy and survey. Comput. Graph. Forum36, 260–284. doi: 10.1111/cgf.13079
6
BoetjeJ.de GraafZ.WopereisI.van GinkelS. O.SmakmanM. H. J.VersendaalJ.et al. (2026). The DIPS model: GenAI integration in digital information problem solving by experts and novices. Comput. Educ.253:105677. doi: 10.1016/j.compedu.2026.105677
7
BråtenI.StrømsøH. I.SalmerónL. (2011). Trust and mistrust when students read multiple information sources about climate change. Learn. Instr.21, 180–192. doi: 10.1016/j.learninstruc.2010.02.002
8
BuçincaZ.MalayaM. B.GajosK. Z. (2021). To trust or to think: cognitive forcing functions can reduce overreliance on AI in AI-assisted decision-making. Proc. ACM Hum.-Comput. Interact.5:188. doi: 10.1145/3449287
9
CaiY.TianS. (2025). Student translators’ web-based vs. GenAI-based information-seeking behavior in translation process: a comparative study. Educ. Inf. Technol.30, 18997–19025. doi: 10.1007/s10639-025-13523-7
10
CarterB. T.LukeS. G. (2020). Best practices in eye tracking research. Int. J. Psychophysiol.155, 49–62. doi: 10.1016/j.ijpsycho.2020.05.010,
11
ChenA.XiangM.ZhouJ.JiaJ.ShangJ.LiX.et al. (2025). Unpacking help-seeking process through multimodal learning analytics: a comparative study of ChatGPT vs human expert. Comput. Educ.226:105198. doi: 10.1016/j.compedu.2024.105198
12
EvansJ. S. B. T.StanovichK. E. (2013). Dual-process theories of higher cognition: advancing the debate. Perspect. Psychol. Sci.8, 223–241. doi: 10.1177/1745691612460685
13
FauthF.González-MartínezJ. (2021). On the concept of learning transfer for continuous and online training: a literature review. Educ. Sci.11:133. doi: 10.3390/educsci11030133
14
ForsythB. R. (2018). Defining far transfer via thematic similarity. Cogent Psychol.5:1523348. doi: 10.1080/23311908.2018.1523348
15
FuY.WengZ. (2024). Navigating the ethical terrain of AI in education: a systematic review on framing responsible human-centered AI practices. Comput. Educ. Artif. Intell.7:100306. doi: 10.1016/j.caeai.2024.100306
16
García-AlonsoE. M.León-MejíaA. C.Sánchez-CabreroR.Guzmán-OrdazR. (2024). Training and technology acceptance of ChatGPT in university students of social sciences: a netcoincidental analysis. Behav. Sci.14:612. doi: 10.3390/bs14070612,
17
GottschlingS.KammererY. (2021). Readers’ regulation and resolution of a scientific conflict based on differences in source information: an eye-tracking study. Discourse Process.58, 468–490. doi: 10.1080/0163853X.2021.1930808
18
GottschlingS.KammererY.GerjetsP. (2019). Readers’ processing and use of source information as a function of its usefulness to explain conflicting scientific claims. Discourse Process.56, 429–446. doi: 10.1080/0163853X.2019.1610305
19
GuanZ.-H.LinS. S. J. (2025). When and how learners engage with source information in digital multiple-text reading: effects of task instruction and text trustworthiness from eye-tracking technology. Comput. Educ.236:105362. doi: 10.1016/j.compedu.2025.105362
20
HesselsR. S.NuthmannA.NyströmM.AnderssonR.NiehorsterD. C.HoogeI. T. C. (2025). The fundamentals of eye tracking part 1: the link between theory and research question. Behav. Res. Methods57:16. doi: 10.3758/s13428-024-02544-8,
21
HoogeI. T. C.NyströmM.NiehorsterD. C.AnderssonR.FoulshamT.NuthmannA.et al. (2026). The fundamentals of eye tracking part 6: working with areas of interest. Behav. Res. Methods58:65. doi: 10.3758/s13428-025-02937-3,
22
JärveläS.MalmbergJ.HaatajaE.SobocinskiM.KirschnerP. A. (2021). What multimodal data can tell us about the students’ regulation of their learning process?Learn. Instr.72:101203. doi: 10.1016/j.learninstruc.2019.04.004,
23
KaiserC.KaiserJ.SchallnerR.SchneiderS. (2025) A new era of online search? A large-scale study of user behavior and personal preferences during practical search tasks with generative AI versus traditional search enginesProceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing SystemsNew York, NY: Association for Computing Machinery
24
KangH. (2021). Sample size determination and power analysis using the G*power software. J. Educ. Eval. Health Prof.18:17. doi: 10.3352/jeehp.2021.18.17,
25
KasneciE.SesslerK.KüchemannS.BannertM.DementievaD.FischerF.et al. (2023). ChatGPT for good? On opportunities and challenges of large language models for education. Learn. Individ. Differ.103:102274. doi: 10.1016/j.lindif.2023.102274
26
LakensD. (2022). Sample size justification. Collabra Psychol.8:33267. doi: 10.1525/collabra.33267
27
LeeH.P.SarkarA.TankelevitchL.DrososI.RintelS.BanksR.et al (2025) The impact of generative AI on critical thinking: self-reported reductions in cognitive effort and confidence effects from a survey of knowledge workersProceedings of the 2025 CHI Conference on Human Factors in Computing SystemsNew York, NY: Association for Computing Machinery
28
LeeJ. D.SeeK. A. (2004). Trust in automation: designing for appropriate reliance. Hum. Factors46, 50–80. doi: 10.1518/hfes.46.1.50_30392,
29
LimC.-Y.InJ. (2019). Randomization in clinical studies. Korean J. Anesthesiol.72, 221–232. doi: 10.4097/kja.19049,
30
LiuX.CuiY. (2025). Eye tracking technology for examining cognitive processes in education: a systematic review. Comput. Educ.229:105263. doi: 10.1016/j.compedu.2025.105263
31
LongD.MagerkoB. (2020) What is AI literacy? Competencies and design considerationsProceedings of the 2020 CHI Conference on Human Factors in Computing SystemsNew York, NY: Association for Computing Machinery
32
Martín-MoncunillD.Alonso MartínezD. (2025). Students’ trust in AI and their verification strategies: a case study at Camilo José Cela University. Educ. Sci.15:1307. doi: 10.3390/educsci15101307
33
MelumadS.YunJ. H. (2025). Experimental evidence of the effects of large language models versus web search on depth of learning. PNAS Nexus4:pgaf316. doi: 10.1093/pnasnexus/pgaf316,
34
MiaoF.HolmesW. (2023). Guidance for Generative AI in Education and Research. Paris: UNESCO.
35
MohammadiM.TajikE.Martinez-MaldonadoR.SadiqS.TomaszewskiW.KhosraviH. (2025). Artificial intelligence in multimodal learning analytics: a systematic literature review. Comput. Educ. Artif. Intell.8:100426. doi: 10.1016/j.caeai.2025.100426
36
OECD (2026). OECD Digital Education Outlook 2026: Exploring Effective Uses of Generative AI in Education. Paris: OECD Publishing.
37
OrquinJ. L.HolmqvistK. (2018). Threats to the validity of eye-movement research in psychology. Behav. Res. Methods50, 1645–1656. doi: 10.3758/s13428-017-0998-z,
38
QianW.YangF.CaoY.YiL.GuR.WangZ. (2026). Generative AI and cognitive load in education: a systematic review of WoS/SSCI-indexed studies through the lens of cognitive load theory. Front. Psychol.17:1921504. doi: 10.3389/fpsyg.2026.1921504,
39
SalmerónL.DelgadoP.MasonL. (2020). Using eye-movement modelling examples to improve critical reading of multiple webpages on a conflicting topic. J. Comput. Assist. Learn.36, 1038–1051. doi: 10.1111/jcal.12458
40
ShneidermanB. (2020). Human-centered artificial intelligence: reliable, safe & trustworthy. Int. J. Hum.-Comput Interact.36, 495–504. doi: 10.1080/10447318.2020.1741118,
41
StadlerM.BannertM.SailerM. (2024). Cognitive ease at a cost: LLMs reduce mental effort but compromise depth in student scientific inquiry. Comput. Hum. Behav.160:108386. doi: 10.1016/j.chb.2024.108386
42
SunX.TianH. (2026). AI-enhanced skill assessment in higher vocational education: a systematic review and meta-analysis. Informatics13:20. doi: 10.3390/informatics13020020
43
TaberK. S. (2018). The use of Cronbach’s alpha when developing and reporting research instruments in science education. Res. Sci. Educ.48, 1273–1296. doi: 10.1007/s11165-016-9602-2
44
TsaiM.-J.WuA.-H.BråtenI.WangC.-Y. (2022). What do critical reading strategies look like? Eye-tracking and lag sequential analysis reveal attention to data and reasoning when reading conflicting information. Comput. Educ.187:104544. doi: 10.1016/j.compedu.2022.104544
45
VasconcelosH.JörkeM.Grunde-McLaughlinM.GerstenbergT.BernsteinM. S.KrishnaR. (2023). Explanations can reduce overreliance on AI systems during decision-making. Proc. ACM Hum.-Comput. Interact.7:129. doi: 10.1145/3579605
46
WalravenA.Brand-GruwelS.BoshuizenH. P. A. (2009). How students evaluate information and sources when searching the world wide web for information. Comput. Educ.52, 234–246. doi: 10.1016/j.compedu.2008.08.003
47
YanL.GreiffS.LodgeJ. M.GaševićD. (2025). Distinguishing performance gains from learning when using generative AI. Nat. Rev. Psychol.4, 435–436. doi: 10.1038/s44159-025-00467-5
48
YanL.GreiffS.TeuberZ.GaševićD. (2024). Promises and challenges of generative artificial intelligence for human learning. Nat. Hum. Behav.8, 1839–1850. doi: 10.1038/s41562-024-02004-5,
49
YenR.SultanumN.ZhaoJ. (2024) To search or to gen? Exploring the synergy between generative AI and web search in programming. In: Proceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems (CHI EA’24), New York, NY: Association for Computing Machinery
Keywords
eye tracking, generative artificial intelligence, immediate near transfer, information seeking, responsible AI, source verification
Citation
Ma J, Yue H and Li X (2026) Responsible use of generative AI in Chinese vocational education: eye-tracking evidence on source verification and immediate near-transfer across three information-seeking conditions. Front. Psychol. 17:1946196. doi: 10.3389/fpsyg.2026.1946196
Received
23 July 2026
Revised
10 September 2026
Accepted
14 September 2026
Published
30 September 2026
Volume
17 - 2026
Updates
Copyright
© 2026 Ma, Yue and Li.
This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.
*Correspondence: Xinyang Li, 51280125079@stu.ecnu.edu.cn
Disclaimer
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.
来源:Frontiers in Psychology · frontiersin.org
猜你喜欢
- 研究:AI 迎合式回应经元认知惰性与依赖降低学习者自主性Frontiers in Psychology · 1 小时前
- 12周课外多元训练对印尼青少年运动能力、认知能力与问题性网络使用的影响:一项随机对照试验Frontiers in Psychology · 1 天前
- 数字健康干预对冠心病患者生活质量、焦虑与抑郁疗效的网络元分析Frontiers in Psychiatry · 1 天前
- Epic Cosmos 230万人电子病历研究:孤独症谱系障碍人群自杀未遂风险分布Frontiers in Psychiatry · 1 天前
- 研究用眼动、EEG 与语义差异量表考察 AI 生成中国水墨画的观看反应Frontiers in Psychology · 1 天前