Study design
This was a retrospective longitudinal diagnostic accuracy study designed to evaluate the performance of a multimodal large language model (LLM) in longitudinal risk stratification and trajectory assessment of oral lichen planus (OLP). The study assessed concordance between LLM-based classification and expert panel consensus using serial clinical documentation and follow-up data. Reporting adhered to the STARD-AI (Standards for Reporting Diagnostic Accuracy Studies for Artificial Intelligence) guidelines. The pre-specified primary endpoint was sensitivity for the detection of expert-defined high-risk cases (one-versus-rest). Secondary endpoints included overall trajectory classification accuracy and three-level risk stratification performance.
Study setting
This retrospective diagnostic accuracy study was conducted using archived clinical records from three academic institutions: the Faculty of Dentistry, King Salman International University (El-Tor, Egypt), the Faculty of Dentistry, Galala University (Suez, Egypt), and the Faculty of Dentistry, Ain Shams University (Cairo, Egypt). Eligible cases were identified through a systematic review of electronic health records and institutional clinical archives at these centers. All data were de-identified before extraction and analysis. Case selection was performed consecutively according to predefined inclusion and exclusion criteria.
Ethics approval and consent to participate
The study protocol was reviewed and approved by the Institutional Review Board of King Salman International University (approval number: IRB014-2025). The requirement for informed consent was waived by the same Institutional Review Board due to the retrospective design and use of fully de-identified archival clinical data, in line with applicable national regulations governing retrospective research. The study was conducted in accordance with the Declaration of Helsinki.
Informed consent for expert participation
All the experts participating in the reference standard assessment process were well informed regarding the objectives and the concept of the study. The experts participated in the study voluntarily and gave their consent to contribute their evaluations for the research and publication. There is no disclosure of personal details related to the experts in this manuscript.
Participants and case selection
Cases were retrospectively identified from institutional electronic health records and clinical archives at the participating centers. Consecutive cases meeting predefined eligibility criteria were screened. Participating centers followed comparable clinical documentation protocols, including standardized intraoral photography and structured clinical records, enabling consistent longitudinal case file construction.
Inclusion criteria
Participants were eligible if they:
-
1.
Were aged ≥ 18 years.
-
2.
Had a confirmed diagnosis of OLP based on modified World Health Organization (WHO) clinical and histopathological criteria.
-
3.
Had complete baseline documentation at the index visit (T0), including standardized intraoral photographs, structured clinical notes describing lesion morphology and symptoms, and histopathological confirmation.
-
4.
A minimum follow-up period of 24 months was selected because clinically meaningful trajectory changes and potential malignant transformation in oral lichen planus often occur over extended surveillance periods. Previous longitudinal OLP studies have reported follow-up durations of approximately 4–5 years or longer when evaluating disease progression and malignant transformation15,16.
For each eligible case, a structured longitudinal case file was constructed incorporating baseline (T0) documentation and all available follow-up clinical notes, interval intraoral photographs, and histopathology reports where applicable. These longitudinal records were used for both index test evaluation and reference standard determination.
Exclusion criteria
Cases were excluded if they:
-
Represented oral lichenoid lesions with identifiable reversible etiologic factors.
-
Had incomplete baseline or follow-up documentation.
-
Were lost to follow-up before 24 months.
-
Had concurrent autoimmune mucosal disorders that could confound progression assessment.
Sample size calculation
The study was powered on the clinically critical endpoint of high-risk progression detection. Pilot data (n = 23) demonstrated a sensitivity of 0.75 and a high-risk prevalence of 17.4%. To estimate sensitivity with a predefined 95% confidence interval precision (half-width) of ± 0.12, at least 52 outcome-positive cases were required. Based on the observed prevalence, this corresponded to a minimum total sample size of 287 patients. To ensure adequate representation across prognostic strata and compensate for potential exclusions, recruitment continued until 300 consecutive eligible patients were included.
Cases with missing baseline or follow-up information were excluded (n = 25). No imputation of missing data was performed; only studies with complete case files that met the inclusion criteria were used.
Reference standard: expert panel prognostic stratification
Trajectory and risk classifications were assigned retrospectively based on the complete longitudinal course rather than solely on baseline (T0) features. Each panelist independently reviewed the complete longitudinal clinical record for every case, including baseline documentation (T0), serial follow-up clinical notes, interval intraoral photographs, and histopathology reports when available. Panelists were blinded to the LLM outputs and to each other’s initial assessments. Based on a comprehensive evaluation of disease evolution over the minimum 24-month follow-up period, each expert assigned:
-
One trajectory classification (stable benign, inflammatory progression, or suspicious malignant evolution), and.
-
One risk category (low, moderate, or high).
The expert panel consisted of three specialists in oral medicine with over 10 years of experience diagnosing and treating oral potentially malignant disorders and oral lichen planus. Before conducting the research, the members of the panel were provided with standardized written definitions of all risk categories and all trajectories. Before independent assessment, panelists also assessed representative example cases to ensure consistent application of the predefined classification criteria. Each panelist independently assigned trajectory and risk classifications. Final reference labels were determined by majority agreement (at least two of three concordant ratings).
The use of majority agreement (two out of three agreements) as the reference standard was predefined because this approach would reduce the influence of inter-individual variation among observers without diminishing the importance of the experts’ subjective evaluations. This method is commonly used when evaluating diagnostic accuracy when there is no specific objective reference standard.
In the rare instances where all three panelists assigned different categories, a structured consensus discussion was conducted to reach a final adjudicated classification. Inter-expert agreement before adjudication was quantified using Fleiss’ kappa.
The inter-expert agreement before adjudication was high (Fleiss κ = 0.79, 95% CI: 0.74–0.84). A complete lack of agreement was seen in 11 cases (3.7%), where a structured discussion was used to reach a consensus. Histopathologically documented cases of dysplasia were seen in 41 of the 66 high-risk cases (62.1%), while the remainder were classified based on suspicious features that warranted biopsy (non-healing ulcers, progressive induration, or exophytic lesions).
Trajectory categories were predefined. Stable benign indicated no documented morphological progression or dysplastic change during follow-up. Inflammatory progression indicated clinical worsening (e.g., expansion of erosive/atrophic areas or increased symptom severity) without dysplasia. Suspicious malignant evolution indicated persistent non-healing ulceration, progressive induration, exophytic change, or histopathological dysplasia. Risk levels reflected overall longitudinal behavior: low (stable, no dysplasia), moderate (persistent inflammatory activity without dysplasia), and high (histopathological dysplasia or clinically suspicious progression warranting biopsy).
Incorporation of histopathologic results was made to form just one element of the longitudinal disease chart, and not an entirely separate gold standard at each follow-up point. Repeat biopsies at each surveillance visit cannot be done routinely for OLP because it would neither be ethical nor practical for stable lesions. Furthermore, trajectory group assignment in OLP is based on the overall development over time of the clinical symptoms and signs, along with histopathology. Thus, the final reference group assignments were based on the overall longitudinal trajectory classification.
Mitigation of potential information leakage
To minimize the risk of information leakage from histopathological reports, these reports were preprocessed before the longitudinal case files were constructed. Specifically, the explicit histopathological information concerning the presence or absence of “dysplasia” and its different grades, such as “dysplastic,” “mild,” “moderate,” or “severe dysplasia,” was removed from the text. No changes were made to the clinical reports, intraoral photographs, or the data.
The use of this masking technique was meant to prevent any shortcut learning that would enable a multimodal model to perform well through the identification of isolated keywords used for diagnosis, instead of drawing insights from clinical and imaging data over time17,18. The objective was to assess whether the model’s classifications were based on information from serial records rather than isolated diagnostic markers.
The masking method was used only on the input data file for the algorithm. The expert panelists were provided access to the unaltered clinical notes, photographs, and pathology reports in the creation of reference standard classifications.
This preprocessing step was carried out using a standard text processing script by one researcher (F.E.A.H.), and this was verified by another researcher (S.M.S.). The rationale behind this process was to minimize the chances that the predictions made by the model were based on isolated high-signal keywords and that it was forced to rely on combined longitudinal clinical and imaging information.
It is noteworthy that although explicit diagnostic terminology was removed, histopathological information is retained, and that, therefore, implicit information cannot be completely excluded.
Exact prompt provided
A single predefined prompt was developed before commencement of the study and was used for all cases without modification. Alternative prompts were not evaluated, and no prompt optimization or iterative prompt engineering was undertaken. This approach was intentionally adopted to avoid performance inflation associated with prompt selection and to provide a standardized and reproducible evaluation framework. Consequently, all cases were assessed using identical instructions and input structures.
“You are an oral medicine specialist. A longitudinal case file will be provided to you for a patient who has a confirmed diagnosis of oral lichen planus. It will contain baseline clinical notes, follow-up notes, and possibly photographs and histopathology reports. Based on this complete longitudinal information provided to you, classify this case as:
-
(1)
Trajectory category (select one): stable benign/inflammatory progression / suspicious malignant evolution.
-
(2)
Risk category (select one): low/moderate/high.
Give a brief explanation for your classification. Do not enter any extra information outside of this classification and explanation.”
Index test: multimodal large language model framework
The index test consisted of ChatGPT-5.2 (OpenAI, San Francisco, CA, USA), a multimodal large language model capable of processing both textual and image-based inputs. ChatGPT-5.2 was accessed through the ChatGPT web interface during the study analysis period. No API-level customization, fine-tuning, or system-level prompt modification was performed. All cases were evaluated using a predefined fixed prompt. A single predefined prompt was developed before commencement of model evaluation and remained unchanged throughout the study. No prompt optimization, iterative refinement, prompt selection, or response regeneration procedures were performed. Clinical narratives, serial intraoral photographs, and histopathology reports were presented chronologically within a single conversation to preserve the temporal sequence of disease evolution. Because the model was accessed through the web interface rather than an application programming interface (API), technical parameters such as temperature settings, seed values, and backend model configurations were not user-configurable.
Case file preparation
For each included patient, a standardized digital case file was constructed incorporating:
-
Baseline clinical documentation (T0), including structured clinical narrative and intraoral photographs.
-
Serial follow-up clinical notes documenting symptom evolution and morphological changes.
-
Follow-up intraoral photographs.
-
Histopathology reports when applicable.
All data were de-identified before submission to the model.
Model implementation
The longitudinal case file was submitted to the LLM using a predefined fixed prompt. Clinical narratives, serial intraoral photographs, and histopathology reports were presented chronologically within a single conversation to preserve the temporal sequence of disease evolution. The same prompt structure and data presentation format were used for all cases.
Longitudinal data structuring and trajectory assessment
For each patient, all available clinical information was organized chronologically from baseline (T0) to the most recent follow-up visit. The longitudinal case file included baseline clinical findings, serial clinical notes, intraoral photographs obtained during follow-up visits, and histopathology reports when available. Information from each visit was presented sequentially to preserve the temporal order of disease evolution. No explicit weighting of individual timepoints was applied. All available visits were presented to the model in chronological order using the same standardized format. Consequently, trajectory assessment relied on the model’s interpretation of patterns across the complete longitudinal record rather than on investigator-defined weighting schemes. Progression was defined as a sustained increase in clinical severity, lesion extent, symptom burden, development of concerning clinical features, worsening histopathological findings, or movement toward malignant transformation risk during follow-up. In contrast, fluctuation referred to temporary changes in disease activity followed by stabilization or return toward the patient’s previous clinical status. The model was instructed to evaluate the overall longitudinal pattern across visits rather than isolated findings at individual timepoints.
Repeatability assessment
For assessing the stability of model output in a deterministic environment, each case was run three times independently. These runs were conducted on separate days to account for possible backend variability. For each of these runs, the main classification for each case in terms of trajectory and risk category was noted. Agreement for all three runs was calculated by using Fleiss’ kappa and percent agreement. For each of those runs, in case of disagreement, the majority classification was used. Cases that had three different outputs, if any, were noted separately.
Assessment of longitudinal information utilization and textual leakage
To evaluate whether model classifications reflected utilization of longitudinal information rather than reliance on isolated diagnostic terminology, a randomly selected subset of 50 cases (16.7% of the total dataset) was analyzed. For each case, the model-generated explanation was independently reviewed by two investigators (F.E.A.H. and S.M.S.) using a predefined binary scoring framework. Explanations were classified as demonstrating longitudinal information utilization if they contained explicit references to temporal changes between visits, disease evolution, progression, regression, or comparisons between baseline and follow-up findings. Explanations lacking such temporal references were classified as not demonstrating longitudinal information utilization. Inter-rater agreement was assessed using Cohen’s kappa statistic. Disagreements were resolved through discussion with a third investigator (A.A.B.).
The cases were considered to demonstrate use of longitudinal information if the rationale made specific reference to time changes, such as “enlargement of erosive area compared to previous visit,” rather than just using quotes of histopathology vocabulary. Consensus was reached by a third party, A.A.B., in the event of disagreements. All 50 cases showed the presence of at least one specific time reference, supporting the conclusion that the model utilized longitudinal data rather than relying on individual diagnostic vocabulary. No cases were excluded in this analysis, as the initial evaluation was performed on the entire analysis.
Statistical analysis
Statistical analyses were performed using IBM SPSS Statistics (Version 27.0.1; IBM Corp., Armonk, NY, USA) and Python. Diagnostic performance for trajectory and risk stratification tasks was evaluated using overall accuracy, sensitivity (recall), specificity, precision (positive predictive value), negative predictive value, and F1-score. Exact (Clopper–Pearson) 95% confidence intervals were calculated for accuracy, sensitivity, and specificity. For the primary endpoint (high-risk detection), a one-versus-rest approach was applied to derive true positive, false negative, false positive, and true negative counts. Agreement between the LLM classifications and expert consensus was assessed using Cohen’s kappa and quadratically weighted kappa for ordinal outcomes. Full confusion matrices were generated, and misclassification patterns, including directionality (under- versus over-classification), were analyzed descriptively. All analyses were two-sided, with a significance threshold of α = 0.05.

