跳到正文
原文
Frontiers in Psychology· Tingting Kuai·· 2 小时前AI 评分30

数字参与如何影响中国传统音乐的文化认同:社会认知传播机制研究

Social cognitive communication mechanism of digital participation in Chinese traditional music in the context of cultural identity and intangible cultural heritage

AI 导读

一项发表于 Frontiers in Psychology 的研究基于 2,346 条音乐、短视频与社交媒体平台行为记录及 1,128 份古琴、昆曲、琵琶与地方戏曲用户问卷,提出数字参与(DP)经社会认知传播(SCC)影响文化认同(CI)的四层架构与阶段过程模型。

正文

Abstract

Digital platforms have reshaped how traditional Chinese music circulates, yet the psychological mechanism linking user participation to cultural identification remains undertheorized. Drawing together social cognitive theory, contemporary cultural psychology, and communication theory, this study develops a four-layer architecture (input, participation, cognitive, and identity) coupled with a staged process model spanning information encoding, symbolic decoding, cognitive reconstruction, and identity internalization. A hybrid specification formalizes the mechanism, pairing structural equation modeling with attention-weighted neural aggregation; both components are estimated and validated on the same data. Empirical testing draws on 2,346 behavioral records from music, short-form video, and social media platforms, paired with 1,128 survey responses from users engaged with the guqin, Kunqu, pipa, and regional opera repertoires. Item loadings, heterotrait–monotrait ratios, collinearity diagnostics, and a common-method check document measurement quality. Results indicate that digital participation (DP) shapes cultural identity (CI) mainly through social cognitive communication (SCC), with the co-creation and dissemination strata showing indirect effects that are more than three times those of passive exposure. As the direct participation-identity path remains significant, the pattern is one of partial rather than full mediation. Cultural identity also looks more like a hinge than a terminal state: a chained specification linking identification to downstream cognitive engagement returns a reliable indirect path, which is consistent with the recursive reading the framework proposes but, on single-wave evidence, cannot establish it. Community climate exerts the strongest moderating influence, exceeding both platform affordance and content polish. Since the design is cross-sectional, all of these patterns are read as associations rather than as causal effects; even so, they suggest that audience cultivation and co-creation tooling deserve at least as much investment as interface refinement in the digitization of traditional music.

1 Introduction

The accelerated diffusion of digital infrastructure has reshaped how intangible cultural heritage (ICH) circulates, archives itself, and reaches new audiences, opening pathways once confined to physical apprenticeship and regional festivals (Stefano and Davis, 2017). Short-video platforms, immersive media, and algorithmic recommendation now allow heritage forms to surface in everyday feeds. Still, the same mechanisms compress nuance, fragment narrative continuity, and expose fragile traditions to decontextualization ().

Within this transformed landscape, Chinese traditional music—encompassing folk song, instrumental traditions, opera, and ritual repertoires—occupies a position of unusual fragility and cultural weight, since its meaning depends as much on embodied performance and communal memory as on sonic content (Zhu, 2024). Recent surveys indicate that although online exposure to guqin, Kunqu, and regional ballads has grown markedly, sustained engagement and intergenerational transmission remain uneven, with younger audiences often treating these forms as aesthetic curiosities rather than living practice (). That gap between exposure and commitment is, at bottom, a psychological problem: what changes inside a listener between a first encounter and a durable sense of belonging, and under what conditions does the change hold (Yang et al., 2022)?

Scholarly attention to these tensions has expanded along several converging lines. A study on cultural identity (CI)—drawing from the constructivist tradition—treats identification as an ongoing negotiation between symbolic resources and lived experience rather than a static inheritance (Pulis, 2014). Research on digital safeguarding of ICH has examined documentation standards, virtual museums, and participatory archiving, often emphasizing technical feasibility while leaving the audience side under-theorized (). Communication scholars have meanwhile probed how networked publics form around heritage content, identifying loops of imitation, remix, and affective resonance that differ sharply from broadcast-era reception (). A parallel literature on traditional music transmission has documented the role of master–apprentice ties, regional ecosystems, and educational policy in sustaining repertoire (Rice, 1987), while a smaller body of study has begun to apply social cognitive theory—particularly observational learning, self-efficacy, and outcome expectancy—to media-mediated cultural behavior ().

Three more recent developments sharpen these questions. Studies of informal digital learning show that online self-efficacy and perceived social presence jointly regulate how much cognitive effort learners invest (Wu, 2023; ). A study on feedback perception indicates that the interpretive frame a learner brings to a platform signals matters at least as much as the signals themselves (). Contemporary cultural psychology has re-established that cognition, emotion, and motivation are culturally patterned processes rather than neutral substrates onto which culture is later layered ().

What these strands have not yet produced is an integrated account of how digital participation (DP) actually shapes cognition and identity in the heritage context. Existing studies tend to isolate a single facet of the problem—platform affordances, audience attitudes, or transmission outcomes—without modeling the recursive process that links them (Wu, 2024). Three questions consequently remain open. Do participatory acts of different cognitive depth carry cultural meaning at different rates, or is participation a single undifferentiated quantity? Does identification, once formed, feed back into the cognitive processing that produced it? And is the translation from participation to identification gated chiefly by technical affordance or by the social climate of the community in which it occurs? Empirical evidence on the mediating role of participatory behaviors—commenting, sharing, remixing, learning—remains sparse and methodologically scattered (), leaving heritage practitioners and platform designers to improvise, with few evidence-based principles to guide content strategy, interaction design, or community cultivation.

These three questions set the agenda for what follows, and they are psychological before they are technological. With digital media now woven into daily routines of listening, learning, and self-presentation, the mechanism at issue concerns observational learning, efficacy beliefs, motivational uptake, and the developmental study of identity formation—processes that unfold inside individuals even when their traces are recorded as platform metrics (). Building on social cognitive theory, contemporary cultural psychology, and communication theory, the study develops a multidimensional model linking platform affordances, user participation behaviors, cognitive and affective processing, and cultural identification, then tests that model against survey and behavioral data drawn from heritage-oriented digital communities ().

Three contributions follow, each stated as a claim that the data can refute rather than as a synthesis of the literature. The first is that participation is stratified in its cultural consequences: acts differing in cognitive investment—watching, reacting, remaking, redistributing—should transmit to identity at sharply different rates, which converts a familiar descriptive typology into a testable ordering. The second is that identity may be re-entrant rather than terminal. Once identification takes hold, it should begin to organize the very cognitive processing that produced it, a proposition an additive framework cannot express because they close the system at outcome measurement. Whether that loop actually runs in time is a question the present design can pose but not settle; what a single wave of data can show is whether the association such a loop implies is present at all. The third is that the translation is socially rather than technically gated, so community climate should outweigh interface affordance in shaping the participation-to-identity path—a prediction that runs counter to the technocentric emphasis of much digital safeguarding work.

Taken together, these claims extend social cognitive theory into a domain where the modeled behavior is cultural rather than instrumental, and they return to cultural psychology a measurable account of how mediated interaction reshapes belonging. The remainder of the article sets out the theoretical foundations, model construction, empirical strategy, results, and their wider implications.

2 Theoretical foundations and research framework

2.1 Cultural identity and the communication of intangible cultural heritage

Cultural identity, as a theoretical construct, has migrated across disciplines over the past half-century, moving from essentialist accounts that treated it as a stable inheritance to relational and processual formulations that read it as continually rewritten through symbolic interaction (). Hall () articulation of identity as both “being” and “becoming” remains a productive anchor, since it allows researchers to hold together the sedimented layers of tradition and the fluid moments of self-positioning that characterize contemporary cultural life. For heritage scholarship, this dual character matters: ICH is at once an inheritance to be safeguarded and a resource to be enacted, and the boundary between preservation and reinvention is rarely as clean as policy documents imply (Smith, 2006).

Contemporary cultural psychology adds a constraint that heritage scholarship has been slow to absorb: culture and mind are mutually constituted, so cultural materials do not merely furnish content for cognition but organize the perceptual, emotional, and motivational processes through which content is registered at all (). On that reading, an encounter with heritage music is never a neutral intake of sound; it is patterned by schemas the culture has already installed.

Within the ICH context, identity construction unfolds along several intersecting dimensions—cognitive recognition of heritage symbols, affective attachment to communal memory, behavioral participation in practice, and reflexive narration of belonging. Chinese traditional music carries these layers with unusual density. A guqin piece is not only a sonic object; it indexes literati aesthetics, regional lineage, and ritualized listening habits that have accumulated over centuries (Yung, 1997). When such repertoire circulates, listeners encounter compressed cultural memory whose interpretation depends on prior exposure, social cues, and the symbolic frames offered by mediating institutions (Sterne, 2009). Prior cultural exposure, therefore, functions as an interpretive resource rather than as a demographic control, since it determines which features of a performance become perceptually available in the first place. Heritage-experience research reports this pattern precisely, with prior involvement conditioning the depth of the mental experience through which identification is assembled (Yang et al., 2022).

To formalize this layered relationship, the cultural identity strength of an individual i toward a heritage form can be expressed as a weighted aggregation of dimensional components; a standard composite formulation in identity research is expressed as follows:

where Dik denotes the score of individual i on identity dimension k (cognitive, affective, behavioral, and narrative), and wk the dimension weight derived from empirical loading.

The link between identity and heritage communication is itself dynamic. Each communicative episode revises the symbolic stock available for future identification, producing what is best described as a recursive loop, and digital environments plausibly accelerate it. Adapting the classic diffusion logic to this recursive case, the temporal evolution of aggregate identification within a digitally mediated audience may be written as follows (Equation 2):

where E(t) captures the intensity of digital exposure and participation, α the assimilation rate, β the decay rate under attentional competition, and CImax the saturation ceiling, following standard adoption-decay dynamics. The formulation makes explicit why platform-driven engagement, rather than mere content availability, has become the pivotal variable in how traditional music sustains identification in everyday digital life (Van Dijck et al., 2018). It should be read as an analytic statement of the dynamics the mechanism implies, not as a model fitted here: a cross-sectional design identifies only the stationary counterpart of such a system, a restriction to which Section 3.3 returns.

2.2 Social cognitive theory and the mechanism of digital participation

Bandura's social cognitive theory advances a triadic reciprocal model in which personal factors, behavior, and environmental conditions co-determine one another, departing from earlier stimulus–response accounts by granting human agency a constitutive role (). Translated into digital settings, the “environment” is no longer a neutral backdrop but an algorithmically curated feed whose recommendations, visibility cues, and interaction affordances continuously reshape what users notice, attempt, and internalize (). Four mechanisms inherited from the original theory acquire new texture under such conditions: observational learning thrives on the abundance of demonstration videos and live streams; symbolic modeling intensifies through avatar-mediated reenactment; visible feedback metrics recalibrate self-efficacy beliefs; and outcome expectancies are formed not only through anticipated personal benefit but through perceived social currency within networked communities ().

Recent evidence from digital learning environments supports the reading with some precision: online self-efficacy and perceived social presence jointly predict the effort learners invest, and the influence of platform feedback depends on how learners construe it rather than on how often it arrives (Wu, 2023; ; ). Heritage participation differs in content but not in structure, which is what licenses the transfer of the theory to this setting.

The triadic interaction can be written in dynamic form, following Bandura's standard formulation, as follows (Equation 3):

where P, B, and E denote personal cognition, behavior, and environment at time t, and f a coupling function capturing reciprocal influence.

Digital participation is best treated as a layered construct rather than a flat behavioral variable. Four ascending strata can be distinguished—exposure (L1), interaction (L2), co-creation (L3), and dissemination (L4)—each demanding different cognitive investment and yielding different feedback signals (van Dijck, 2009). A composite participation index DPi for user i aggregates these strata with empirically derived weights as expressed in the following (Equation 4):

where Lij is the standardized score of user i at participation stratum j and λj the weight reflecting that stratum's cognitive demand, derived through confirmatory factor loading in the manner standard for formative composite measurement ().

Coupling these layers to social cognitive formation requires a mediating mechanism rather than a direct path. The specification adopted here holds that digital participation acts on social cognition, SC, through self-efficacy, SE, and outcome expectancy, OE, with environmental affordance A moderating the gain, as expressed in the following (Equation 5):

Following the moderated-mediation specification widely adopted in social cognitive research (). The expression makes explicit that exposure alone is insufficient; cognitive uptake depends on whether participatory acts generate efficacy gains and credible expectancy signals within an enabling platform environment (Schunk and Usher, 2019). Motivationally, this is the difference between a platform that shows users what others have made and one that lets them discover what they can make themselves; only the second generates the mastery information that informs revisions to efficacy beliefs (). This framework supplies the analytical scaffolding for the model construction that follows.

2.3 Research framework and variable definition

Drawing the two preceding strands together, this study specifies a research framework organized around three core constructs: digital participation (DP), social cognitive communication (SCC), and cultural identity (CI). Digital participation is treated as the antecedent driver, social cognitive communication as the mediating mechanism, and cultural identity as the distal outcome, with platform affordance entering as a contextual moderator (). The framework rests on the logic that participatory acts in heritage-oriented digital communities generate cognitive and affective traces that, in turn, rework the symbolic resources from which identification is assembled (). What distinguishes the specification from earlier participation-identity models is not the presence of a mediator, which is by now conventional, but the treatment of participation as ordered by cognitive depth and of identity as potentially re-entrant. Both restrictions are stated sharply enough to fail, though only the first can be decisively tested on data gathered at a single point in time.

Operationally, DP retains the four-stratum specification introduced earlier (exposure, interaction, co-creation, and dissemination). SCC is decomposed into three observable facets: cognitive elaboration (CE), affective resonance (AR), and normative endorsement (NE), each measured through validated multi-item scales adapted to the heritage music setting (). CI is captured through cognitive recognition, affective attachment, behavioral commitment, and narrative articulation, consistent with the composite expression in Equation 1. Item wordings were adapted from published instruments and re-anchored to the heritage music setting; the full item pool, together with the translation and back-translation record, is described in Section 4.1.

The mediating role of SCC between DP and CI is formalized as follows:

where Xim denotes control covariates (age, prior musical training, and platform tenure), and ϵi, νi are stochastic disturbances. The mediating effect is the product γ1θ2, following Baron and Kenny's classical decomposition adapted for structural equation modeling (SEM) estimation ().

The following is an expression of a moderated path that captures the conditioning role of affordance A:

The interaction term SCCi×Ai tests whether platforms offering richer participatory affordances amplify the translation of cognitive engagement into identification, following standard moderation specification ().

From this structure, four hypotheses follow. H1: DP positively predicts CI. H2: SCC mediates the DP−CI relation. H3: Among the four participation strata, co-creation (L3) and dissemination (L4) exert stronger indirect effects on CI than exposure (L1). H4: Platform affordance A moderates the SCC−CI link, with stronger effects under high-affordance conditions. These hypotheses are tested in the empirical analysis reported in Section 4.

3 Construction of the social cognitive communication mechanism model for digital participation in Chinese traditional music

3.1 Overall architecture design of the digital participation communication mechanism

The architecture developed here joins social cognitive theory's triadic logic with the processual reading of cultural identity, organizing the mechanism into four hierarchically nested layers—input, participation, cognitive, and identity—threaded by a moderating affordance band and a closing feedback loop (). Each layer carries a distinct functional load while remaining recursively coupled to its neighbors, so that communication of Chinese traditional music in digital settings is treated as a stratified flow rather than a linear transmission (Poell et al., 2019).

The input layer aggregates the symbolic and technical antecedents that seed the mechanism: heritage repertoire (guqin pieces, opera arias, and regional ballads), digitization standards, creator outputs, and the recommender logic that governs visibility. These inputs do not act directly on users; they pass through participation affordances that decide which sonic objects become candidates for engagement (). The participation layer then channels users into the four ascending strata introduced earlier—exposure, interaction, co-creation, and dissemination—producing a behavioral substrate captured by the composite index DPi in Equation 4.

What gives the architecture its analytical bite is the cognitive layer above participation. Behavioral traces are translated into cognitive elaboration, affective resonance, and normative endorsement, with self-efficacy and outcome expectancy serving as internal calibrators that filter which experiences gain symbolic weight (Schunk and DiBenedetto, 2021). The identity layer, finally, accumulates these cognitive deposits into the four-dimensional identification profile specified in Equation 1, feeding back into participation through revised motivation and selective attention. Within the architecture, this loop is constitutive rather than decorative: identification is modeled as a continuously updated disposition that biases subsequent participatory choices (). The status of that claim warrants careful consideration. It belongs, for the moment, to the model rather than to the evidence, since cross-sectional data can reveal the association such a loop implies but cannot observe the loop unfolding. Read psychologically, the proposal is identity formation in the developmental sense—exploration through participation, commitment through repeated confirmation, reconsideration when the community's signals change ().

Figure 1 maps the architecture and its principal pathways.

Figure 1

Table 1 presents the layered structure, specifying constituent elements, functional roles, and key variables that anchor the subsequent measurement.

Table 1

LayerCore elementsFunctional roleKey variables
InputHeritage repertoire, digital archives, creator outputs, and recommender algorithmsSupply symbolic and technical antecedentsContent richness and algorithmic visibility
ParticipationExposure, interaction, co-creation, and disseminationConvert exposure into behavioral substrateL1-L4 scores and DPi index
CognitiveCognitive elaboration, affective resonance, and normative endorsementTranslate behavior into symbolic uptakeCE, AR, NE; SE, and OE
IdentityRecognition, attachment, commitment, and narrationAggregate cognition into identificationCIi composite and weights wk
Affordance bandInterface design, community norms, and visibility cuesModerate cross-layer translation gainAffordance index A
Feedback loopRevised motivation and algorithmic re-rankingRe-seed input with updated signalsLoop gain and decay rate β

Hierarchical structure and functional positioning of the digital participation communication mechanism.

Two design choices distinguish the architecture from earlier accounts. Treating affordance as a cross-cutting moderator rather than a layer-specific attribute lets the model accommodate platform heterogeneity without inflating its parametric burden. Reading the feedback loop as a constitutive element, rather than closing the system at the level of outcome measurement, captures the cumulative drift of repeated digital encounters that descriptive heritage studies have noted but rarely formalized (). Figure 1 should therefore be read as an estimation map rather than as an illustration. Its forward arrows correspond one-to-one with parameters recovered below: the participation-to-cognition translation is the first-stage coefficient of Equation 6, the cognition-to-identity translation the second-stage coefficient of Equation 7, the affordance band the interaction term of Equation 8, and the feedback loop the chained specification examined—though not temporally identified—in Section 4.3. This stratified scaffolding underwrites the variable specification and structural estimation that follow in Sections 3.2 and 3.3.

3.2 Multidimensional digital participation–social cognition coupling communication process

Having fixed the layered architecture in Section 3.1, the argument now turns to the process side: how a user actually moves over time from a first encounter with a digitized piece of traditional music to a stabilized act of cultural identification. The flow unfolds across four coupled stages—information encoding, symbolic decoding, cognitive reconstruction, and identity internalization—each shaped by an interplay of platform, content form, community interaction, and algorithmic recommendation ().

Encoding begins when creators, archivists, or heritage institutions translate sonic and performative material into platform-native formats: short videos, livestream fragments, immersive audio, interactive scores. Format choice is far from neutral; a 15-s Douyin clip and a 30-min Bilibili lecture encode the same guqin piece into radically different cognitive openings (). Decoding follows as the recommender surfaces the encoded artifact to a user whose interpretive resources—prior exposure, musical training, peer cues—determine whether the symbolic payload registers as ornament, curiosity, or invitation (Schäfer, 2011). Community interaction at this point matters more than is usually granted: a thread of informed comments can reframe an opaque passage into a culturally legible one within seconds.

Cognitive reconstruction sits at the pivot of the process. Decoded fragments are stitched into the user's existing schemas through elaboration, affective tagging, and normative comparison with what peers are seen to value (). The dynamics of this stitching can be written, following standard reinforcement formulations, as follows (Equation 9):

where Ci(t) is the cognitive state of the user i at time t, Ri(t) the reconstruction signal generated by the current participatory episode, and η a learning rate sensitive to affordance richness and emotional salience (Sutton and Barto, 2018). Internalization, the final stage, occurs when reconstructed cognition consolidates into a disposition stable enough to bias future selection, sharing, and self-narration—closing the loop back to Section 3.1's feedback channel ().

Figure 2 traces the staged flow and its principal couplings.

Figure 2

Table 2 presents the distinguishing features of each stage to their dominant cognitive outputs.

Table 2

StageDominant driverUser behaviorCognitive outputIdentity trace
Information encodingCreator, archive, and format designPassive reception of pushed contentAttention capture and schema primingLatent recognition cue
Symbolic decodingRecommender + peer cuesBrowsing and light reaction (like and save)Interpretive parsing and salience taggingAwareness fragment
Cognitive reconstructionCommunity interactionCommenting, learning, and remixingElaboration, affective resonanceEmerging attachment
Identity internalizationSelf-reflection and narrative practiceSharing, teaching, and repertoire selectionNormative endorsement and schema consolidationStabilized identification
Recursive feedbackUpdated motivationSelective re-engagementReinforced cognitive biasIdentification accrual

Stage characteristics of digital participation and corresponding cognitive outputs.

Two features of the process deserve emphasis. The transitions are asymmetric: encoding-to-decoding is high-throughput but low-conversion, whereas reconstruction-to-internalization is low-throughput but high-conversion, which means platform strategies aimed solely at exposure volume tend to leave the most consequential transition underserved (). The lateral inputs—platform, content form, community, and algorithm—do not act uniformly across stages; community weight rises sharply at reconstruction, while algorithmic gating dominates the earlier transitions. Figure 2 is likewise tied to the estimation rather than free-standing. Its four stages are not measured as separate time points; they enter the empirical model as the salience weights recovered by the attention block and reported in Section 4.3, where the pivotal status claimed for reconstruction becomes a quantity that can be checked rather than asserted. This staged, asymmetric, laterally conditioned flow is what the structural model in Section 3.3 will operationalize.

3.3 Mathematical modeling and algorithmic implementation of core mechanism elements

The staged process specified in Section 3.2 invites a more compact mathematical treatment, one that can be estimated against survey and behavioral data without losing the layered structure of the mechanism. The specification draws on structural equation modeling for the latent-variable backbone and cognitive diffusion dynamics for the temporal extensions, and concludes with an attention-weighted aggregation suited to heterogeneous participation traces (). As a hybrid formalism invites the suspicion that some of its parts are decorative, the division of labor is worth stating at the outset: the latent-variable block carries the hypothesis tests, the attention block is estimated and validated against held-out data in Section 4.3, and the continuous-time expressions define the dynamic system of which the estimated cross-sectional model is a snapshot. Which of these are fitted to data, and which remain analytic scaffolding, is made explicit at the close of this section.

The digital participation intensity for user i at time t is specified as a saturating function of stratum-weighted behavioral inputs as expressed as follows (Equation 10):

with κ controlling diminishing returns, following standard saturation specification (). The following logistic-type equation governs the diffusion of social cognition across an interacting user population (Equation 11):

where μ is the assimilation coefficient, ρ the attentional decay, and SCmax the cognitive ceiling. Peer-driven contagion is layered on through a network term as expressed in the following (Equation 12):

with aik the tie strength between users i and k, following standard network-diffusion form ().

Cultural identity convergence is then captured by an exponential approach to a moving target shaped by sustained cognitive input as expressed in the following (Equation 13):

with τ the convergence rate. The following expression signifies the heterogeneous participatory signals that are fused through a multilayer perceptron (MLP)–style hidden representation:

where xi stacks the stratum scores, W1and W2 are weight matrices, b1and b2 are biases, and σ the rectified linear unit (ReLU) activation in the standard MLP formulation (). Attention weighting then re-scores stages by salience as expressed in the following:

producing the weighted propagation effect provided as follows:

which feeds the outcome layer through the following equation:

For the latent-variable backbone, parameter estimation follows maximum likelihood under the SEM measurement specification, which is expressed as follows (Equation 18):

with Λ the loading matrix, Φ the latent covariance, and Θδ the residual covariance, in classical Jöreskog form (). End-to-end training of the neural pathway minimizes a regularized loss and expressed as follows:

optimized through an adaptive stochastic gradient method with early stopping on a held-out fold (). Table 3 inventories the core variables and symbols for clarity before estimation.

Table 3

SymbolMeaningType
DPi(t)Digital participation intensity for user iComposite
LijStratum score (exposure–dissemination)Observed
λjStratum weightParameter
SCi(t)Social cognition stateLatent
μ, ρAssimilation, decay coefficientsParameter
aikPeer tie strengthNetwork input
CIi(t)Cultural identity convergenceOutcome
αisAttention weight on stage sLearned
ΠiWeighted propagation effectDerived
ΘδSEM residual covarianceParameter

Core model variables and symbol specification.

This hybrid specification keeps the SEM block answerable to theory while letting the attention-weighted block absorb non-linearities a purely linear path would otherwise push into the residual. Two clarifications guard against overreading it. The components actually fitted to data are the measurement and structural equations, the mediation and moderation decompositions, and the attention-weighted aggregation, whose implementation, training protocol, and out-of-sample performance are reported in Section 4.3. The continuous-time expressions for diffusion, contagion, and convergence are analytic scaffolding: they state the dynamic system the mechanism implies, but a cross-sectional design identifies only its stationary counterpart, so their parameters are not estimated here. The recursive reading of identity belongs to that scaffolding as well. What Section 4.3 tests is the association the loop would produce, not the loop itself, and this restriction is explicitly carried forward into the limitations stated in Section 6.

4 Empirical analysis and discussion of results

4.1 Data collection and sample characteristics

Empirical material for testing the mechanism was assembled from two complementary streams between March and October 2025. The behavioral stream draws on publicly observable user traces from three categories of platforms—digital music services (NetEase Cloud Music and QQ Music), short-video platforms (Douyin, Kuaishou, and Bilibili), and social media (Weibo and Xiaohongshu). Candidate accounts were identified through platform search and hashtag crawling using a fixed keyword list covering guqin, Kunqu, pipa, erhu, and regional opera repertoires. They were retained only if they had interacted at least three times with content carrying those tags during the observation window. That threshold was set to exclude one-off algorithmic encounters while remaining low enough to keep genuinely passive viewers inside the frame (). Sampling was purposive rather than probabilistic, so the resulting sample represents engaged heritage-music audiences on Chinese platforms rather than the general population.

The survey stream complements these traces with a structured questionnaire administered to identifiable users who consented to participate, capturing the latent constructs—self-efficacy, outcome expectancy, cognitive and affective processing, and the four identity dimensions—that behavioral logs cannot reach. Invitations were issued via platform direct message to all 2,346 accounts in the behavioral pool, and 1,128 usable responses were returned, an effective response rate of 48.1%. All items were measured on seven-point Likert scales. Participation items were adapted from established measures of user-generated content engagement, efficacy and expectancy items from social cognitive instruments, and identity items from established cultural identification scales, each re-anchored to the heritage music setting. English source items were translated into Chinese and independently back-translated by a bilingual researcher who had not seen the originals; discrepancies were resolved through discussion before piloting on 60 respondents, whose data are excluded from the analytic sample.

After deduplication, bot filtering, and removal of inactive accounts, 2,346 valid behavioral records were retained, of which 1,128 were paired with completed questionnaires through a hashed account identifier. Preprocessing followed a four-step pipeline: log cleaning, semantic tagging of content categories, standardization of stratum-level participation scores, and missing-value imputation using expectation-maximization for items with less than 5% missingness ().

Two elements of that pipeline deserve fuller description, since they carry the majority of the interpretive weight. Semantic tagging was performed by two trained coders working independently from a codebook that assigned each observed act to one of the four participation strata and each content item to a repertoire category. The coders overlapped on a random 20% of the corpus. They returned chance-corrected agreement of 0.87 for stratum assignment and 0.84 for repertoire category, with remaining disagreements adjudicated by the third author.

Stratum scores were then standardized within platform before aggregation, so that platform-specific base rates of liking and commenting would not be mistaken for differences in participation depth, and the stratum weights entering the composite index were fixed by confirmatory factor loading rather than assigned a priori. Analyses were conducted in R with lavaan for the structural model and in Python with PyTorch for the attention block.

Both data streams involved human participants, and the ethics arrangements differed between them. The Research Ethics Committee of Jiangsu Vocational Institute of Commerce reviewed and approved the full protocol (Reference Number: JSIC-IRB-2025-0317). Questionnaire respondents provided electronic informed consent on the instrument's opening page, before any item was displayed; the consent text outlined the purpose of the study, the voluntary nature of participation, the right to withdraw at any point without consequence, and the anonymous storage of responses. For the behavioral stream, the committee approved a waiver of individual consent, on the grounds that the traces were publicly posted, non-sensitive, and collected in line with the platforms' terms of service. Account identifiers were hashed at the time of collection; no private messages or direct identifiers were retained; and all results are reported in aggregate, so that no individual user can be identified from the subsequent analyses.

Reliability and validity checks returned satisfactory indicators. Cronbach's α for the participation, cognition, and identity scales ranged from 0.84 to 0.91, and composite reliability ranged from 0.86 to 0.93. Convergent validity, assessed through average variance extracted, exceeded the 0.50 threshold across all latent constructs, while discriminant validity was confirmed under the Fornell–Larcker criterion ().

Table 4 summarizes the sample's demographic spread and stratum-level participation. Two patterns stand out: younger users dominate the active strata, and formal musical training correlates with markedly higher composite participation, suggesting that pre-existing cultural capital still gates depth of engagement even in algorithmically open environments.

Table 4

AttributeCategoryn%Mean DPi
GenderFemale/male1,302/1,04455.5/44.50.58/0.54
Age18–2589438.10.62
Age26–3576532.60.59
Age36–5048720.80.49
Age>502008.50.41
TrainingFormal musical background61226.10.67
Tenure>3 years on platform1,48863.40.61
StratumCo-creators (L3 active)53822.90.74

Sample demographic profile and digital participation distribution.

As Figure 3 illustrates, demographic composition varies sharply by platform type; Figure 4 makes clear that the strata of participation are unevenly distributed, a heterogeneity that the structural analysis in Section 4.2 must absorb ().

Figure 3

Figure 4

Because a structural model is only as trustworthy as the measurement model beneath it, the measurement properties are reported in full before any path is interpreted. Table 5 gives, for each latent construct, the number of retained items, the range of standardized loadings, composite reliability, average variance extracted, and the largest variance inflation factor among its indicators. Every loading clears 0.65 and is significant at the 0.001 level; composite reliability ranges from 0.86 to 0.93; average variance extracted exceeds 0.50 throughout; and no variance inflation factor approaches the conventional threshold of 3.3, so multicollinearity does not distort the coefficients reported later. The confirmatory measurement model fits acceptably, with a chi-square-to-degrees-of-freedom ratio of 2.18, Root Mean Square Error of Approximation (RMSEA) of 0.045, Standardized Root Mean Square Residual (SRMR) of 0.041, Comparative Fit Index (CFI) of 0.957, and Tucker–Lewis Index (TLI) of 0.949.

Table 5

ConstructItemsLoading rangeCRAVEMax VIF
Digital participation (DP)120.682–0.8410.890.542.31
Cognitive elaboration (CE)50.714–0.8620.900.642.07
Affective resonance (AR)50.706–0.8490.880.601.94
Normative endorsement (NE)40.691–0.8330.860.611.88
Self-efficacy (SE)40.735–0.8710.910.682.15
Outcome expectancy (OE)40.702–0.8480.890.632.02
Cultural identity (CI)120.674–0.8790.930.582.44
Platform affordance (A)50.688–0.8270.870.571.76
Content quality40.699–0.8360.860.601.69
Community climate40.721–0.8580.900.661.83

Measurement model assessment: item loadings, reliability, convergent validity, and collinearity diagnostics.

AVE, average variance extracted; Max VIF, variance inflation factor.

Discriminant validity was assessed twice, since the Fornell–Larcker criterion alone is known to miss violations that the heterotrait–monotrait ratio detects (; Voorhees et al., 2016). Table 6 reports the heterotrait–monotrait ratios above the diagonal, the square root of average variance extracted on the diagonal, and the inter-construct correlations below it. No ratio exceeds 0.72, comfortably under the 0.85 benchmark, and each diagonal entry exceeds the correlations in its row and column, so the constructs are empirically distinct.

Table 6

ConstructDPCEARNECIA
DP0.7350.5710.5240.4980.5310.437
CE0.5120.8000.6720.6210.6590.451
AR0.4680.6010.7750.6480.6270.418
NE0.4410.5570.5830.7810.6120.403
CI0.4760.5940.5610.5480.7620.412
A0.3890.4020.3710.3580.3650.755

Discriminant validity: heterotrait–monotrait ratios (above diagonal), square root of average variance extracted (diagonal), and construct correlations (below diagonal).

Common method bias was addressed procedurally and statistically. Behavioral and self-report sources were kept separate, item order was randomized, and anonymity was assured at the point of administration. A marker-variable test using a theoretically unrelated construct left the substantive path coefficients essentially unchanged, and the first unrotated factor accounted for 28.4% of the variance, well-below the level at which method variance would be a plausible rival explanation (Podsakoff et al., 2024). With measurement quality established on these several fronts, the structural estimates that follow can be read as relations among constructs rather than as artifacts of the instrument.

4.2 Empirical test of digital participation effects on social cognitive communication

Building on the measurement model validated above, the structural model specified in Equations 6–8 was estimated on the paired sample of 1,128 respondents using maximum likelihood with Satorra-Bentler robust standard errors, after confirming acceptable model fit (χ2/df = 2.31, RMSEA = 0.048, CFI = 0.951, and TLI = 0.942) under the cutoffs recommended for SEM applications in communication research (). Direct and indirect pathways were decomposed using the standard mediation expression provided as follows (Equation 20):

with bias-corrected bootstrap confidence intervals based on 5,000 resamples. Multi-group analysis was then run across three contrasts—participation stratum, music type (instrumental, vocal/opera, and ritual), and age cohort—using the chi-square difference test on nested invariance models ().

The full-sample estimates anchor the picture. The direct effect of DP on SCC returned γ1 = 0.482 (p < 0.001); SCC on CI returned θ2 = 0.413 (p < 0.001); and the residual direct DP→CI path stood at θ1 = 0.187. As the direct path remains significant, indirect effect of 0.199, alongside a total effect of 0.386, accounts for roughly 52% of the total effect, indicating partial rather than full mediation. In the terminology of Zhao and colleagues, direct and indirect effects share a sign, so the pattern is more precisely described as complementary mediation, and the label is used consistently in this sense throughout Sections 4.2 and 4.3 (Zhao et al., 2010). The moderating interaction SCC×A produced θ3 = 0.156 (p = 0.002), supporting H4. Effect sizes were summarized through the following expression (Equation 21):

yielding f2 = 0.21 for the mediator block, a medium effect under Cohen's convention ().

Stratum decomposition delivered the more revealing contrast. Indirect effects through SCC rose sharply from exposure (L1: 0.071) and interaction (L2: 0.142) to co-creation (L3: 0.246) and dissemination (L4: 0.231), confirming H3. The pairwise difference between L3 and L1 was tested using the following expression (Equation 22):

returning z = 4.86 (p < 0.001).

The music-type comparison, presented in Table 7, shows that instrumental repertoires (guqin, pipa, and erhu) yield stronger indirect transmission than vocal and opera forms, a gap that plausibly reflects the lower interpretive entry barrier for sonically immersive instrumental clips on short-video feeds. Age-group analysis gave the inverse pattern, with the cohort comparison ratio (Equation 23):

indicating that younger users channel participation into cognition with substantially greater efficiency, though their baseline identification level remains lower than that of older cohorts.

Table 7

PathEstimateSEt-valueDecision
DP→CI (direct)0.1870.0414.56***H1 supported
DP→SCC0.4820.03812.68***–
SCC→CI0.4130.0459.18***–
DP→SCC→CI0.1990.024–H2 supported
L3 indirect via SCC0.2460.0298.48***H3 supported
L1 indirect via SCC0.0710.0223.23**H3 supported
SCC×A→CI0.1560.0503.12**H4 supported
Instrumental music0.2210.0346.50***–
Vocal/opera music0.1780.0394.56***–

Path coefficient estimates and hypothesis test results.

Read together, Figures 5, 6 support the hierarchical logic anticipated in Section 3.3: depth of participation, not breadth, tracks cognitive uptake, and platform affordance appears to amplify the conversion rather than to substitute for it ().

Figure 5

Figure 6

4.3 Mediation and moderation analysis of the cultural identity construction mechanism

The mediation logic posited in Section 2.3 is here subjected to a more demanding test. The model was re-specified with cultural identity treated alternately as an endogenous outcome and as an internal mediator nested within the SCC→CI pathway, then examined how platform type, content quality, and community climate condition the transmission. Bootstrap mediation followed the standard product-of-coefficients approach with 5,000 resamples (Preacher and Hayes, 2008), and conditional indirect effects were derived through (Equation 24):

with M denoting the moderator and γ3, θ4 the corresponding interaction coefficients (). The Johnson–Neyman technique was applied to locate moderator regions of significance (Equation 25):

where Δ is the discriminant of the conditional effect equation, following Hayes' formulation ().

The mediation block returned a stable indirect effect of DP on CI through SCC at 0.199 (95% CI [0.151, 0.247]). Since the direct path retains significance, this is again partial, complementary mediation rather than the full mediation that a non-significant direct effect would indicate (Zhao et al., 2010). A second specification then treated cultural identity itself as an internal mediator between cognitive elaboration and behavioral commitment; the chained indirect path yielded an estimate of 0.124 (95% CI [0.089, 0.163]). That estimate is consistent with the re-entrant role sketched in Section 2.3, yet it is worth being blunt about what it cannot do. All variables in the chain were measured at a single point in time, so the ordering comes from the model rather than from the data, and cross-sectional estimates of mediated effects are known to diverge substantially from their longitudinal counterparts (). The same coefficient would arise if users who already identify strongly simply elaborate more, which is why the result is reported here as an association compatible with recursion rather than as evidence of it.

Moderation results sharpen the picture. Platform type interacted with SCC at θ4 = 0.183 (p < 0.001), with high-affordance platforms (Bilibili and Xiaohongshu) yielding conditional indirect effects roughly 1.7 times those of low-affordance feeds. Content quality, indexed by production polish and informational depth, moderated the first-stage path at γ3 = 0.142 (p = 0.004). Community climate, measured through reciprocity and supportive tone, produced the largest second-stage gain at θ4 = 0.211 (p < 0.001), with the effect size expressed as Equation 26:

returning across the three moderators jointly (). Robustness was checked against alternative model specifications, with consistency index:

across K = 6 subsamples, indicating stable parameter recovery.

Table 8 makes the conditioning pattern legible: community climate registers the largest moderating weight, exceedingly even platform affordance—a ranking that runs counter to the technocentric expectation and that the remainder of the analysis therefore probes rather than assumes.

Table 8

Path/effectEstimate95% CIpDecision
DP→SCC→CI (partial, complementary mediation)0.199[0.151, 0.247]< 0.001Supported
DP→CE→CI→BC (chained)0.124[0.089, 0.163]< 0.001Supported
Platform type moderation0.183[0.108, 0.258]< 0.001Supported
Content quality moderation (1st-stage)0.142[0.045, 0.239]0.004Supported
Community climate moderation (2nd-stage)0.211[0.137, 0.285]< 0.001Supported
High-affordance conditional indirect0.276[0.214, 0.338]< 0.001Stronger
Low-affordance conditional indirect0.162[0.108, 0.216]< 0.001Weaker
High-quality content × DP0.318[0.241, 0.395]< 0.001Stronger
Warm community climate × SCC0.347[0.268, 0.426]< 0.001Strongest
Robustness consistency ρrobust0.078––Stable

Mediation and moderation test results comparison.

Two notes keep the table honest. Its first row reports partial, complementary mediation, not full mediation: the direct DP→CI coefficient in Table 7 remains significant at 0.187, the total effect is 0.386, and the indirect effect accounts for roughly 52% of it. Its final row reports the mean relative deviation of parameter estimates across six subsamples, so the value of 0.078 falls below the 0.10 tolerance set in Equation 27 rather than above the fit threshold.

Figures 7, 8 together carry the theoretical payoff. The mechanism is robust, but its conversion efficiency is socially rather than technically gated: where communities cultivate reciprocity, and content carries interpretive depth, the participation-to-identification translation accelerates, regardless of platform polish.

Figure 7

Figure 8

A hybrid model earns its second component only by demonstrating that the component adds something, so the attention-weighted block specified in Equations 14–17 was implemented and validated rather than left as a formal gesture. Stratum scores, platform indicators, and stage-level behavioral aggregates were stacked as inputs to a two-layer perceptron with 16 hidden units and rectified linear activation; the attention layer scored the four process stages, and the weighted representation was mapped to the identity outcome through a linear read-out. Training minimized the regularized objective of Equation 19 under Adam, with the paired sample split into five folds, an inner validation split for early stopping, and regularization strength chosen by grid search on that inner split alone, so no test fold informed model selection. Figure 9 reports the comparison of interest.

Figure 9

Two things follow. Out-of-sample explained variance rises from a mean of 0.411 for the structural model alone to 0.462 for the augmented specification, a modest but consistent gain of roughly five percentage points that holds across all folds; the attention block therefore captures non-linearity the linear path leaves in the residuals without displacing the theoretical model. The learned weights, moreover, are far from uniform: cognitive reconstruction absorbs 0.351 of the total saliences against 0.164 for information encoding. That the two blocks, estimated under different assumptions, agree on where the mechanism concentrates is the strongest internal check the design affords.

5 Discussion

The empirical pattern that emerged across Sections 4.2 and 4.3 supports a reading that goes beyond confirming the hypothesized paths. Three interlocking observations organize what follows; the findings are then set against prior study and against the alternative explanations this design cannot rule out.

First, digital participation appears to drive social cognitive transmission not through the volume of exposure but through the depth to which observational learning, symbolic modeling, and community feedback interweave. The sharp rise in the indirect effect from the exposure stratum to co-creation, from 0.071 to 0.246, suggests that watching a Kunqu clip and remixing one are categorically different acts. The first leaves a transient perceptual trace; the second engages symbolic modeling in Bandura's strict sense, since the user must select, reorganize, and reproduce elements of the heritage form, generating mastery information and outcome feedback that exposure alone never supplies.

This is where the psychological reading earns its keep: enactive experience of remaking a phrase is precisely the source from which efficacy beliefs are revised, and the digital-learning literature documents the same asymmetry between passive reception and productive effort (Wu, 2023; ). Community interaction during co-creation then supplies the corrective and reinforcing cues that turn imitation into internalized practice—a result sitting awkwardly with the still-common assumption that platform reach is itself a proxy for cultural transmission.

Second, cultural identity looks more like a hinge than an endpoint. The chained mediation estimate of 0.124 is consistent with a reading in which identification not only accrues from cognitive engagement but also comes to organize it, filtering which guqin pieces a user pauses on, which comments they trust, which lineages they begin to claim. Consistency, however, is not sequence. A single wave of data cannot separate that reading from its mirror image, in which prior identification simply prompts deeper elaboration, and the methodological literature warns that cross-sectional estimates of such chains can be badly misleading about the longitudinal process they are taken to represent (). Read through contemporary cultural psychology, the proposal is that culture and mind mutually constitute one another on short time scales: cultural material supplies schemas, schemas organize attention, and reorganized attention selects further cultural material (). It remains, on this evidence, a proposal.

It also implies that identity here takes a context-bound form—among younger Douyin users, an aestheticized self-presentation; among Bilibili learners, a quasi-apprentice belonging; among older Weibo audiences, a custodial stance. What differs across these subgroups is less the strength of identification than its form, which is the sense in which digital interaction transforms identity rather than merely increasing it, and which earlier study tended to flatten by treating identification as a single magnitude (), (Yang et al., 2022). As the subgroup contrasts rest on descriptive comparison rather than on tests of measurement invariance, they are offered as an interpretive proposal awaiting confirmation.

Third, the shaping logic of platforms, algorithms, and user agency is neither one-way nor symmetrical. Algorithms front-load attention but cannot manufacture conviction; platform affordances raise the ceiling of possible engagement, then yield to community climate, which in these data carried the heaviest moderating weight. User agency, often theorized in the abstract, surfaces here as a negotiation between algorithmic offering and self-directed search—a negotiation more visible among trained users, who navigate against the grain of recommendation to reach repertoire they value.

Comparison with prior accounts produces both convergence and friction. The mediation pattern is consistent with existing findings on networked heritage publics and with heritage-experience research in which involvement reaches identification only through an intervening experiential state rather than directly (Yang et al., 2022). The evidence that community climate outweighs platform affordance, by contrast, diverges from the technocentric tilt of earlier digital safeguarding literature, which has tended to treat interface capability as the binding constraint. The age-cohort inversion—younger users converting participation into cognition more efficiently while maintaining lower baseline identification—complicates the generational pessimism pervading policy discourse on traditional music and aligns with developmental studies in which adolescence and young adulthood are marked by active exploration rather than settled commitment ().

Several alternative explanations cannot be excluded, and they temper the readings above. Selection is the most serious: users who already identify strongly with traditional music may seek out co-creation rather than be changed by it, thereby reproducing the observed ordering with the causal arrow reversed. Algorithmic endogeneity is a second issue, since recommender systems allocate exposure non-randomly, so measured participation depth partly reflects what the platform chose to show. Third, identity and efficacy were both self-reported; the marker-variable check found no evidence of serious inflation, yet shared method variance cannot be eliminated by design (Podsakoff et al., 2024). Fourth, community climate was measured perceptually, so users disposed toward identification may also perceive their communities as warmer. A cross-sectional design cannot adjudicate among these accounts. Longitudinal panels, or experiments in which the depth of participation is manipulated before identification is measured, are the appropriate remedies, and the present results are best read as specifying what such designs should test.

Three risks deserve frank acknowledgment. Cognitive bias accumulates when algorithmic filtering narrows exposure to already-popular repertoire, leaving less circulated forms (ritual music, minority traditions) under-rehearsed in the public imaginary. Cultural misreading intensifies when short-form encoding strips contextual scaffolding, turning a guqin meditation into background ambiance. Homogenization creeps in as creators optimize for platform legibility, smoothing regional and lineage-specific texture into a generic “traditional” aesthetic. None of these is inherent to digital mediation; each is a tractable design problem.

In practice, the findings point toward prioritizing co-creation affordances over polished consumption interfaces, cultivating community moderation that rewards interpretive depth rather than reach, and pairing algorithmic recommendation with curated long-form companions that supply the contextual frames that short clips cannot carry. These are priorities calibrated to associations observed within a single cultural ecosystem, not to general design laws. What the evidence supports for heritage institutions working in comparable settings is a reordering in which audience cultivation deserves attention alongside content digitization rather than after it.

6 Conclusion

This study set out to clarify how digital participation in Chinese traditional music translates into cultural identification through a social-cognitive communication mechanism, and to subject that mechanism to empirical scrutiny. The study delivered a four-layer architecture (input, participation, cognitive, and identity) coupled with a staged process model (encoding, decoding, reconstruction, and internalization), formalized through hybrid SEM and attention-weighted neural specifications, with both components estimated and validated on a paired behavioral–survey dataset of 2,346 records spanning music, short-video, and social media platforms.

Three core findings consolidate the contribution. Depth of participation, rather than volume, is what tracks cognitive transmission, with the co-creation and dissemination strata showing indirect effects on identity more than three times those of passive exposure; because the direct path survives, the mediation is partial rather than full. Cultural identity is better read as a hinge than as a terminal outcome, and the chained estimate is consistent with identification feeding back into cognitive processing—a reading that the present design supports as a hypothesis for longitudinal study rather than as an established sequence. Community climate carries the heaviest moderating weight, exceeding platform affordance and content polish, which invites a reordering of the priority assumptions still common in heritage digitization policy.

Theoretically, the study integrates social cognitive theory, cultural identity theory, and contemporary communication frameworks into a single, tractable architecture, providing a measurable mechanism where earlier literature offered only parallel descriptions. It also returns something to psychology in exchange: an account of how mediated participation revises efficacy beliefs, motivational uptake, and the form—not merely the strength—of cultural belonging. Methodologically, the hybrid specification, whose neural component is validated against held-out data rather than asserted, provides a route to treating heritage transmission as both a latent variable and a learning process. Practically, the findings favor investment in community cultivation, co-creation tooling, and curated long-form companions that anchor the contextual meaning short clips cannot carry. These implications are offered as priorities warranted by the present evidence within Chinese-language platform ecologies rather than as universal prescriptions; claims sometimes made for heritage media in the broader language of cultural influence reach well past what a single cross-sectional study can support.

Limitations should be stated plainly, and several bear directly on how much weight the findings can carry. The design is cross-sectional, so the mediation and chained-mediation results describe covariation consistent with the proposed ordering rather than causal sequence; the recursive role attributed to identity is the claim most exposed here, and a simulation study shows that single-wave estimates of mediated effects can depart sharply from the longitudinal parameters they are meant to approximate (). Cross-lagged panel designs, or experiments that manipulate participation depth before identification is measured, are the appropriate remedies, and the present estimates are best treated as specifying what such studies should target. Sampling was purposive and confined to Chinese-language platforms and predominantly mainland users. Hence, the findings speak to a single cultural ecosystem, and to engaged audiences within it, rather than to a general population. The latent constructs rest on self-report and share a common method; procedural separation and a marker-variable check found no serious distortion, but residual method variance cannot be excluded by design.

Behavioral logs capture observable acts while underrepresenting the affective texture that ethnographic methods would surface, and the mapping from logged acts to participation depth involves coding judgments that, however reliable, remain interpretive. The dynamic expressions in Section 3.3 are not estimated because cross-sectional data identify only their stationary counterparts. Finally, the mechanism treats algorithmic curation as a moderator rather than as an endogenous agent whose own learning dynamics shape what users encounter.

The future study can extend in three directions. Cross-cultural comparison—pairing Chinese traditional music with comparable ICH cases in Japan, Korea, or Southeast Asia—would test the universality of the affordance vs. community ordering. Multimodal data fusion combining audio features, visual cues, network signals, and physiological response would tighten measurement of cognitive reconstruction, and longitudinal designs tracking identity formation across repeated encounters would speak to the developmental dynamics that cross-sectional evidence can only imply (). Intelligent transmission mechanisms, in which recommender systems are co-designed with heritage criteria rather than engagement metrics alone, mark the most consequential frontier; whether such systems can sustain identification without flattening cultural texture is arguably the open question that should guide the next phase of inquiry.

Statements

Data availability statement

The original contributions presented in the study are included in the article/supplementary material, further inquiries can be directed to the corresponding author/s.

Ethics statement

This study involved human participants. It was approved by the Research Ethics Committee of Jiangsu Vocational Institute of Commerce (Reference Number: JSIC-IRB-2025-0317), whose review covered both the questionnaire survey and the collection of publicly observable platform traces. All questionnaire respondents provided electronic informed consent on the opening page of the instrument before any item was presented; the consent statement described the purpose of the study, the voluntary nature of participation, the right to withdraw at any time without consequence, and the anonymous storage and analysis of responses. For the behavioral stream, the committee approved a waiver of individual informed consent because the traces were publicly posted and non-sensitive and were gathered in accordance with the platforms' terms of service, on condition that account identifiers be hashed at the point of collection and that no private content or direct identifier be retained. All data were analyzed and are reported only in aggregate form. The study was conducted in accordance with the Declaration of Helsinki and applicable national regulations.

Author contributions

TK: Writing – review & editing, Writing – original draft. JQ: Writing – review & editing, Writing – original draft. WX: Writing – original draft, Writing – review & editing.

Funding

The author(s) declared that financial support was received for this work and/or its publication. This study was supported by the 2026 General Project of Philosophy and Social Sciences Research in Higher Education of Jiangsu Province, “Research on the Integration Paths of Local Traditional Cultural ‘Intangible Cultural Heritage IP+' in Education and Teaching of Higher Vocational Colleges” (Project No. 2026SJYB226).

Conflict of interest

The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

Generative AI statement

The author(s) declared that Generative AI was not used in the creation of this manuscript.

Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.

Publisher’s note

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.

Abbreviations

ICH, Intangible Cultural Heritage; SCT, Social Cognitive Theory; SCC, Social Cognitive Communication; DP, Digital Participation; CI, Cultural Identity; CE, Cognitive Elaboration; AR, Affective Resonance; NE, Normative Endorsement; SE, Self-Efficacy; OE, Outcome Expectancy; SEM, Structural Equation Modeling; MLP, Multilayer Perceptron; ML, Maximum Likelihood; RMSEA, Root Mean Square Error of Approximation; CFI, Comparative Fit Index; TLI, Tucker–Lewis Index; AVE, Average Variance Extracted; CR, Composite Reliability; CIE, Conditional Indirect Effect; IRB, Institutional Review Board; HTMT, Heterotrait–Monotrait Ratio; VIF, Variance Inflation Factor; CMB, Common Method Bias; SRMR, Standardized Root Mean Square Residual.

References

  • 1

    AguinisH.EdwardsJ. R.BradleyK. J. (2017). Improving our understanding of moderation and mediation in strategic management research. Organ. Res. Methods20, 665–685. doi: 10.1177/1094428115627498

  • 2

    AikenL. S.WestS. G. (1991). Multiple Regression: Testing and Interpreting Interactions. Thousand Oaks, CA: Sage.

  • 3

    AmaralI.FloresA. M. M.AntunesE. (2024). “Gender across digital platforms,” in Young Adulthood Across Digital Platforms, eds. J. Coffey, A. Dobson and M. M. Gleeson (Bingley: Emerald Publishing), 35–56. doi: 10.1108/978-1-83753-524-820241003

  • 4

    AngelilloM. (2025). “Appropriating intangible cultural heritage,” in Intangible Cultural Heritage and New Methodological Frameworks, eds. E. R. Kosmidou and L. G. McMurtry (London: Routledge), 104–120. doi: 10.4324/9781003415329-10

  • 5

    BanduraA. (2001a). Social cognitive theory of mass communication. Media Psychol.3, 265–299. doi: 10.1207/S1532785XMEP0303_03

  • 6

    BanduraA. (2001b). Social cognitive theory: an agentic perspective. Annu. Rev. Psychol.52, 1–26. doi: 10.1146/annurev.psych.52.1.1

  • 7

    BantaB. M. (2018). Participatory heritage, edited by Henriette Roued-Cunliffe and Andrea Copeland. Arch. Issues39, 61–66. doi: 10.31274/archivalissues.11058

  • 8

    BaoX.YuS. (2018). “Research on new media protection for intangible cultural heritage from the perspective of communication,” in Proceedings international conference contemporary education, social sciences and ecological studies (CESSES) (Moscow). doi: 10.2991/cesses-18.2018.140

  • 9

    BaronR. M.KennyD. A. (1986). The moderator–mediator variable distinction in social psychological research: conceptual, strategic, statistical considerations. J. Pers. Soc. Psychol.51, 1173–1182. doi: 10.1037/0022-3514.51.6.1173

  • 10

    BassF. M. (2004). Comments on ‘A new product growth for model consumer durables: the Bass model'. Manage. Sci.50, 1833–1840. doi: 10.1287/mnsc.1040.0300

  • 11

    BolinG. (2017). Media Generations: Experience, Identity and Mediatised Social Change. London: Routledge.

  • 12

    BonacchiL.BevanA.PettD.Keinan-SchoonbaertA. (2019). Participation in heritage crowdsourcing. Mus. Manag. Curatorsh.34, 166–182. doi: 10.1080/09647775.2018.1559080

  • 13

    BonacchiL.KrzyzanskaA. (2019). Digital heritage research re-theorised: ontologies and epistemologies in a world of big data. Int. J. Herit. Stud.25, 1235–1247. doi: 10.1080/13527258.2019.1578989

  • 14

    BranjeS. (2022). Adolescent identity development in context. Curr. Opin. Psychol.45:101286. doi: 10.1016/j.copsyc.2021.11.006

  • 15

    BranjeS.de MoorE. L.SpitzerJ.BechtA. I. (2021). Dynamics of identity development in adolescence: a decade in review. J. Res. Adolesc.31, 908–927. doi: 10.1111/jora.12678

  • 16

    BucherT. (2018). If… Then: Algorithmic Power and Politics.Oxford: Oxford University Press.

  • 17

    CantorA. R. (2019). “Cultural beliefs and everyday practices,” in Tourism and Maternal Health: Customs, Beliefs, and Everyday Practices, eds. N. B. McDonnell and E. E. Stevens (Lanham, MD: Lexington Books), 65–80. doi: 10.5040/9781978738362.ch-6

  • 18

    CaoH.-Q.HanC.-W. (2024). The effect of Chinese vocational college students' perception of feedback on online learning engagement: academic self-efficacy and test anxiety as mediating variables. Front. Psychol.15:1326746. doi: 10.3389/fpsyg.2024.1326746

  • 19

    CaoY. (2023). “Analysis of Chinese digital music in the context of new media,” in Proceedings 2022 4th international conference literature, art and human development (ICLAHD 2022) (Xi'an), 823–829. doi: 10.2991/978-2-494069-97-8_104

  • 20

    CentolaD. (2018). How Behavior Spreads: The Science of Complex Contagions.Princeton, NJ: Princeton University Press. doi: 10.23943/9781400890095

  • 21

    CheungG. W.Cooper-ThomasH. D.LauR. S.WangL. C. (2024). Reporting reliability, convergent and discriminant validity with structural equation modeling. Asia Pac. J. Manag.41, 745–783. doi: 10.1007/s10490-023-09871-y

  • 22

    DottoriniE.IoannidesM.MenchetelliV. (2026). “Beyond the digital documentation: the monumental cemetery of Perugia and interpretive models for funerary heritage,” in Advances in Digital and Cultural Tourism Management, eds. M. Ioannides, J. Martins and L. Cantoni (Cham: Springer), 145–159. doi: 10.1007/978-3-032-22974-8_11

  • 23

    FalbelD. (2021). “madgrad: ‘MADGRAD' method for stochastic optimization,” in CRAN: Contributed Packages. Vienna: CRAN, R Foundation for Statistical Computing. doi: 10.32614/CRAN.package.madgrad

  • 24

    GayoM. (2025). Classical music audiences: evidence from Chile. Cult. Trends, 1–21. doi: 10.1080/09548963.2025.2540958

  • 25

    GeeJ. P. (2013). The Anti-Education Era: Creating Smarter Students through Digital Learning. New York, NY: Palgrave Macmillan.

  • 26

    GeismarH. (2017). “Instant archives?,” in The Routledge Companion to Digital Ethnography, ed. L. Hjorth, H. Horst, A. Galloway, and G. Bell (London: Routledge), 331–343.

  • 27

    GiaccardiE. (ed.). (2012). Heritage and Social Media: Understanding Heritage in a Participatory Culture. London: Routledge.

  • 28

    GoodfellowI.BengioY.CourvilleA. (2016). Deep Learning. Cambridge, MA: MIT Press.

  • 29

    GuoF. (2023). Research on short video communication of intangible cultural heritage. Front. Soc. Sci. Technol.5:51716. doi: 10.25236/FSST.2023.051716

  • 30

    HairJ. F.RisherJ. J.SarstedtM.RingleC. M. (2019). When to use and how to report the results of PLS-SEM. Euro. Bus. Rev.31, 2–24. doi: 10.1108/EBR-11-2018-0203

  • 31

    HairJ. F.SarstedtM.RingleC. M.GuderganS. P. (2018). Advanced Issues in Partial Least Squares Structural Equation Modeling. Thousand Oaks, CA: Sage.

  • 32

    HallS. (1996). “Who needs identity?,” in Questions of Cultural Identity, eds. S. Hall and P. du Gay (London: Sage), 1–17.

  • 33

    HayesA. F. (2022). Introduction to Mediation, Moderation, and Conditional Process Analysis: A Regression-Based Approach, 3rd Edn. New York, NY: Guilford Press.

  • 34

    HayesA. F.MontoyaA. K. (2017). A tutorial on testing, visualizing, and probing an interaction involving a multicategorical variable in linear regression analysis. Commun. Methods Meas.11, 1–30. doi: 10.1080/19312458.2016.1271116

  • 35

    HayesA. F.RockwoodN. J. (2020). Conditional process analysis: concepts, computation, and advances in the modeling of the contingencies of mechanisms. Am. Behav. Sci. 64, 19–54. doi: 10.1177/0002764219859633

  • 36

    HelbergerN.KarppinenK.D'AcuntoL. (2018). Exposure diversity as a design principle for recommender systems. Info. Commun. Soc.21, 191–207. doi: 10.1080/1369118X.2016.1271900

  • 37

    HenselerJ.RingleC. M.SarstedtM. (2015). A new criterion for assessing discriminant validity in variance-based structural equation modeling. J. Acad. Mark. Sci.43, 115–135. doi: 10.1007/s11747-014-0403-8

  • 38

    HuL.BentlerP. M. (1999). Cutoff criteria for fit indexes in covariance structure analysis: conventional criteria versus new alternatives. Struct. Equ. Modeling Multidiscipl. J.6, 1–55. doi: 10.1080/10705519909540118

  • 39

    JöreskogK. G.OlssonU. H.WallentinF. Y. (2016). Multivariate Analysis with LISREL. Cham: Springer. doi: 10.1007/978-3-319-33153-9

  • 40

    KitayamaS.SalvadorC. E. (2024). Cultural psychology: beyond East and West. Annu. Rev. Psychol.75, 495–526. doi: 10.1146/annurev-psych-021723-063333

  • 41

    KlineR. B. (2023). Principles and Practice of Structural Equation Modeling, 5th Edn. New York, NY: Guilford Press.

  • 42

    LachenbruchP. A. (1989). Statistical power analysis for the behavioral sciences (2nd ed.), by J. Cohen. J. Am. Stat. Assoc.84, 1096–1097. doi: 10.2307/2290095

  • 43

    LaRoseR. (2010). The problem of media habits. Commun. Theory20, 194–222. doi: 10.1111/j.1468-2885.2010.01360.x

  • 44

    LittleR. J. A.RubinD. B. (2019). Statistical Analysis with Missing Data, 3rd Edn. Hoboken, NJ: Wiley. doi: 10.1002/9781119482260

  • 45

    LiuY. (2026). Empirical study on the communication effect of textile intangible cultural heritage based on digital storytelling and short video platforms. Text. Leather Rev.9, 3245–3256. doi: 10.31881/TLR.2026.3245

  • 46

    MacdonaldS. (2013). Memorylands: Heritage and Identity in Europe Today. London: Routledge. doi: 10.4324/9780203553336

  • 47

    MaxwellS. E.ColeD. A. (2007). Bias in cross-sectional analyses of longitudinal mediation. Psychol. Methods12, 23–44. doi: 10.1037/1082-989X.12.1.23

  • 48

    O'NeillM. W.TownsendF. C. (eds.). (2002). “Deep foundations 2002: an international perspective on theory, design, construction, and performance,” in Proceedings of the international deep foundations congress 2002 (Orlando, FL).

  • 49

    PanX. (2022). Exploring the multidimensional relationships between educational situation perception, teacher support, online learning engagement, and academic self-efficacy in technology-based language learning. Front. Psychol.13:1000069. doi: 10.3389/fpsyg.2022.1000069

  • 50

    PapacharissiZ. (2015). Affective Publics: Sentiment, Technology, and Politics. Oxford: Oxford University Press.

  • 51

    PodsakoffP. M.PodsakoffN. P.WilliamsL. J.HuangC.YangJ. (2024). Common method bias: it's bad, it's complex, it's widespread, and it's not easy to fix. Annu. Rev. Organ. Psychol. Organ. Behav.11, 17–61. doi: 10.1146/annurev-orgpsych-110721-040030

  • 52

    PoellT.NieborgD.Van DijckJ. (2019). Platformisation. Internet Policy Rev. 8. doi: 10.14763/2019.4.1425

  • 53

    PreacherK. J.HayesA. F. (2008). Asymptotic and resampling strategies for assessing and comparing indirect effects in multiple mediator models. Behav. Res. Methods40, 879–891. doi: 10.3758/BRM.40.3.879

  • 54

    PulisJ. W. (2014). “Religion, diaspora, and cultural identity: an introduction,” in Religion, Diaspora, and Cultural Identity: A Reader in the Anglophone Caribbean, ed. J. W. Pulis (London: Routledge), 17–28. doi: 10.4324/9781315078519

  • 55

    RiceT. (1987). Toward the remodeling of ethnomusicology. Ethnomusicology31:469. doi: 10.2307/851667

  • 56

    SchäferM. (2011). Bastard Culture! How User Participation Transforms Cultural Production.Amsterdam: Amsterdam University Press. doi: 10.5117/9789089642561

  • 57

    SchunkD. H.DiBenedettoM. K. (2021). Self-efficacy and human motivation. Adv. Motiv. Sci.8, 153–179. doi: 10.1016/bs.adms.2020.10.001

  • 58

    SchunkD. H.UsherE. L. (2019). “Social cognitive theory and motivation,” in The Oxford Handbook of Human Motivation, ed. R. M. Ryan (Oxford: Oxford University Press), 9–26. doi: 10.1093/oxfordhb/9780190666453.013.2

  • 59

    SmithL. (2006). Uses of Heritage. London: Routledge

  • 60

    StefanoM. L.DavisP. (2017). The Routledge Companion to Intangible Cultural Heritage. London: Routledge. doi: 10.4324/9781315716404

  • 61

    SterneJ. (2009). “The preservation paradox in digital audio,” in Sound Souvenirs: Audio Technologies, Memory and Cultural Practices, eds. K. Bijsterveld and J. van Dijck (Amsterdam: Amsterdam University Press), 55–65.

  • 62

    SuttonR. S.BartoA. G. (2018). Reinforcement Learning: An Introduction, 2nd Edn. Cambridge, MA: MIT Press.

  • 63

    van DijckJ. (2009). Users like you? Theorizing agency in user-generated content. Media Cult. Soc.31, 41–58. doi: 10.1177/0163443708098245

  • 64

    Van DijckJ.PoellT.de WaalM. (2018). The Platform Society: Public Values in a Connective World. Oxford: Oxford University Press.

  • 65

    VoorheesC. M.BradyM. K.CalantoneR.RamirezE. (2016). Discriminant validity testing in marketing: an analysis, causes for concern, proposed remedies. J. Acad. Mark. Sci.44, 119–134. doi: 10.1007/s11747-015-0455-4

  • 66

    WuR. (2023). The relationship between online learning self-efficacy, informal digital learning of English, and student engagement in online classes: the mediating role of social presence. Front. Psychol14:1266009. doi: 10.3389/fpsyg.2023.1266009

  • 67

    WuS. (2024). Digital transmission of intangible cultural heritage based on empirical modal analysis in a multicultural perspective. Appl. Math. Nonlinear Sci.9, 1–17. doi: 10.2478/amns.2023.2.00629

  • 68

    YangW.ChenQ.HuangX.XieM.GuoQ. (2022). How do aesthetics and tourist involvement influence cultural identity in heritage tourism? The mediating role of mental experience. Front. Psychol.13:990030. doi: 10.3389/fpsyg.2022.990030

  • 69

    YungB. (1997). Celestial Airs of Antiquity: Music of the Seven-String Zither of China. Madison, WI: A-R Editions.

  • 70

    ZhaoX.LynchJ. G.ChenQ. (2010). Reconsidering Baron and Kenny: myths and truths about mediation analysis. J. Consum. Res.37, 197–206. doi: 10.1086/651257

  • 71

    ZhuH. (2024). Research on the new path of inheritance and development of Chinese fine traditional culture in the new era. J. Sociol. Ethnol.6, 115–119. doi: 10.23977/jsoce.2024.060418

Keywords

Chinese traditional music, communication mechanism, cultural identity, digital participation, intangible cultural heritage, social cognitive theory

Citation

Kuai T, Qian J and Xu W (2026) Social cognitive communication mechanism of digital participation in Chinese traditional music in the context of cultural identity and intangible cultural heritage. Front. Psychol. 17:1896221. doi: 10.3389/fpsyg.2026.1896221

Received

31 May 2026

Revised

25 August 2026

Accepted

27 August 2026

Published

07 October 2026

Volume

17 - 2026

Updates

Copyright

© 2026 Kuai, Qian and Xu.

This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.

*Correspondence: Wei Xu, xuweii8108@126.com

Disclaimer

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.

来源:Frontiers in Psychology · frontiersin.org

猜你喜欢