跳到正文
原文
Frontiers in Psychology· Anna Panzeri·· 3 小时前AI 评分24

Frontiers in Psychology 社论:临床与动力心理学中的测量效度推进

Editorial: Advancing measurement validity in clinical and dynamic psychology

AI 导读

Frontiers in Psychology 发表社论指出,临床与动力心理学测量的是无法直接观测的潜变量构念,效度并非测验一次获得便永久持有的证书,而是需随人群、语言、场景与决策情境不断重新积累的可修正证据。

正文

EDITORIAL article

Front. Psychol., 02 October 2026

Sec. Quantitative Psychology and Measurement

Volume 17 - 2026 | https://doi.org/10.3389/fpsyg.2026.1924253

Clinical and dynamic psychology involve objects that cannot be held in the hand. Psychosis risk, anhedonia, dysregulated rage, embodied cognition, and spiritual wellbeing—these are latent constructs inferred from a person's responses rather than read off a dial. A cardiologist can point to a stenosed artery; we can only point to a score and argue about its meaning. That asymmetry carries consequences that the field too easily forgets: in clinical and dynamic psychology, the instrument is the evidence—or, more precisely, the inference we draw from its score is and that inference holds only when a validity argument ties the observed response to a clearly specified population, context, and purpose (; ). When the instrument is miscalibrated, every inference built on it inherits the error: diagnoses misclassify, genuine change goes undetected, and interventions are aimed at something that was never quite there.

This Research Topic takes this idea and puts it to work. Its contributions span new scale development, cross-cultural adaptation, re-examination of established screening tools, and item-level psychometric analysis. They all share the stance that gives this editorial its title. Validity is not a certificate that a test earns once and keeps. Rather, it is a cumulative, revisable body of evidence that must be re-earned whenever the measure moves to a new population, language, setting, decision, or moment in a patient's course. Properly understood, validation never ends. We argue that this never-ending quality is not a nuisance to be wished away but the very form that scientific and clinical responsibility takes when our objects are latent.

1 Validity is never finished

Outside the clinic, the case for measurement rigor is by now familiar. Validity is a unitary, evolving argument about whether scores support the inferences we draw from them, not a property that a test simply possesses (). The replication crisis is, in large part, a measurement crisis: noisy, poorly validated instruments do not merely shrink effects—they can inflate and distort them under selection for statistical significance (). A striking share of psychology's measurement failures trace not to exotic statistics but to a “measurement schmeasurement” attitude, in which scoring and validation decisions go unreported, unexamined, and improvised on the fly (; , ). No sample size and no analytic sophistication can repair an invalid measure.

In clinical and dynamic psychology, the stakes differ in kind, because the person who bears the cost of a bad measurement is not a reader but a patient. The clinimetric tradition emphasizes this point: patient-reported outcomes earn their place through interpretability, responsiveness, and usefulness for the individual case, not through internal homogeneity alone (; ; ; ,). Therefore, what a clinically grounded appraisal asks of any instrument is narrow and demanding at once: not whether it is tidy, but whether it helps the clinician understand the person and decide what to do next ().

2 What the contributions teach

The studies gathered here do not argue this thesis in the abstract; they enact it, each taking responsibility for an instrument at a different stage of its life. Read for their clinical payload rather than their statistics, they fall into three categories.

2.1 Naming what patients feel

Some clinical experiences have outrun the tools meant to capture them. Michałkiewicz et al. give dysregulated anger a careful anatomy, separating its affective and behavioral facets in the Anatomy of Rage Scale; Wiesmann et al. operationalize the intrusive “high-place” urge—the call of the void—that clinicians often hear described but rarely measure; and Xu et al. refine the assessment of anhedonia with a Specific Loss of Interest and Pleasure Scale that retains its meaning from community samples through to patients. What unites these studies is their discipline at the front end: the construct is defined by theory and lived experience before any number is attached to it (). Dickens et al. make the stakes explicit, showing how laypersons actually read the items of the Authentic and Hubristic Pride Scales—because when a respondent and clinician do not share the same understanding of a word, the resulting score measures confusion rather than the patient's state (; ). A high reliability coefficient cannot rescue a measure whose items are read in divergent ways, and asking one narrow question in many ways only flatters the statistic while starving the clinical picture (; ).

2.2 Questioning the frontline screen

Other contributions turn a skeptical eye on instruments already in daily use, where complacency is most dangerous. Sergi et al. reevaluate the factor structure of the Montreal Cognitive Assessment, and D'Ignazio et al. ask whether the Mini-Mental State Examination can truly separate mild cognitive impairment from ordinary aging; neither study treats a familiar screening tool as beyond question. Tolomeo et al. apply Item Response Theory to the Sniffin' Sticks identification subtest, showing how precision in detecting olfactory loss of neurological significance depends on the behavior of individual items. Balhara et al. anchor a screen for problematic internet and device use to an explicit ICD-11 definition rather than to convenience. The shared warning is the one that Fried's work made unavoidable: a single total score can hide wholly different patients behind an identical number; therefore, what appears to be one disorder may in fact be many (; , ; ). What clinical work requires is a screen that grades severity and tracks change, not one that merely correlates ().

2.3 The same score for different people

Finally, several studies confront the quiet assumption that a measure has the same meaning for everyone who completes it. Barajas et al. validate a Spanish CAARMS-S so that psychosis-risk interviews retain their meaning beyond their language of origin; Süto et al. bring a multidimensional inventory of religious and spiritual wellbeing into Hungarian; and Zheng et al. build an embodied-cognition scale for Chinese university students—each treating cultural adaptation as fresh validation rather than mere translation. Kristensen and Vestad emphasize this point, cautioning that some apparent robust sex differences in adolescent mental health may be artifacts of non-equivalent measurement rather than facts about young people. Where comparability fails, the same cut-off silently misclassifies different groups in different ways, and a difference manufactured by the instrument is mistaken for a difference in the patients. Establishing measurement invariance is therefore a precondition for any group comparison, not an optional refinement ().

3 A manifesto for never-ending, patient-centered validation

Taken together, these contributions converge on a single mandate: a measure earns trust only by proving itself where it will actually be used—not on a convenient student sample, displayed thereafter the same way as a trophy. For clinical and dynamic psychology, we distill that mandate into four commitments to never-ending validation (Table 1).

  • Define before you count. Specify the construct from theory and lived experience and confirm that respondents read it as intended before any number is attached.

  • Distrust the familiar screen. Treat widespread use as a reason for scrutiny, not a substitute for it; one total score may conceal distinct patients behind a single number.

  • Establish equivalence before comparing. Adaptation demands new evidence regarding item meaning, response options, structure, and interpretation in the target context—translation is not validation.

  • Measure change, not just consistency. A clinical instrument must move with the patient, remain tolerable to someone who is unwell, and fit the realities of day-to-day clinical practice (; ).

Table 1

No.CommitmentWhat it requiresRisk if ignored
1Define before you countSpecify the construct from theory and lived experience, and confirm respondents read items as intended - before any number is attached.The score measures confusion, not the patient; no reliability coefficient can repair that.
2Distrust the familiar screenTreat widespread use as a reason for scrutiny, not a substitute for it: re-examine structure, cut-offs, and item functioning.One total score hides distinct patients behind a single number - one disorder/construct may be many.
3Establish equivalence before comparingBefore comparing groups, cultures, or sexes, show the measure means the same thing in each. Translation is not validation.The same cut-off misclassifies different groups, and a difference made by the instrument is mistaken for a real one.
4Measure change, not just consistencyA clinical instrument must move with the patient, stay tolerable when they are unwell, and fit the working day of real services.The instrument is deaf to the very change that clinical decisions depend on.

Four commitments of validation: a practical standard for measurement validity in clinical and dynamic psychology.

None of the four is met once and for all; each is renewed as patients deteriorate and recover, as measures cross borders, and as the decisions a score must support keep changing.

This Research Topic collectively enacts these four commitments; each pairs a principle with the clinical failure it prevents, and none is satisfied once and for all.

None of these commitments is met once and for all. Each must be renewed as patients deteriorate and recover, as measures cross borders, and as the decisions we ask a score to support keep changing. That is what never-ending validation means: the day we declare a clinical measure finally, finished, and universally “valid” is the day we stop watching it, and stop protecting the individuals it exists to serve. In clinical and dynamic psychology, validation that never ends is not a burden but the precise shape that care takes when the thing we are measuring is the human mind.

Statements

Author contributions

AP: Conceptualization, Writing – original draft, Writing – review & editing. AK: Conceptualization, Writing – original draft, Writing – review & editing. SM: Conceptualization, Writing – original draft, Writing – review & editing. AAR: Conceptualization, Writing – original draft, Writing – review & editing.

Conflict of interest

The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

Generative AI statement

The author(s) declared that Generative AI was used in the creation of this manuscript. The author(s) declare that Generative AI (Claude, Anthropic) was used in the preparation of this work, exclusively for linguistic refinement and to improve the clarity and readability of the text. The content and scientific concepts presented in this manuscript are entirely the work of the authors. All AI-assisted output was critically reviewed and revised by the authors to ensure scientific accuracy. The authors take full responsibility and accountability for the content of the published work.

Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.

Publisher’s note

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.

References

Keywords

clinical and dynamic psychology, clinimetrics, construct validation, measurement, patient-reported outcomes, psychometrics, questionable measurement practices, validity

Citation

Panzeri A, Klocek A, Mannarini S and Rossi AA (2026) Editorial: Advancing measurement validity in clinical and dynamic psychology. Front. Psychol. 17:1924253. doi: 10.3389/fpsyg.2026.1924253

Received

30 June 2026

Accepted

17 August 2026

Published

02 October 2026

Volume

17 - 2026

Updates

Copyright

© 2026 Panzeri, Klocek, Mannarini and Rossi.

This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.

*Correspondence: Anna Panzeri, anna.panzeri@unibg.it

Disclaimer

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.

来源:Frontiers in Psychology · frontiersin.org

猜你喜欢