More than 5000 MDT switching events were identified across three cancer cohorts
We aimed to identify the cMDT profile of melanoma, ovarian cancer, and AML across the TuPro and TCGA cancer cohorts. In total, we analysed 1117 cancer patient samples: 641 primary and 117 metastatic biopsies. Since no matched normal samples were available in the two cohorts, we downloaded RNA-seq data from 818 skin, ovary, and blood normal samples from GTEx (Fig. 2a). On average, we identified 6983 genes (6756 are protein-coding) in cancer samples and 6033 genes (5887 are protein-coding) in the GTEx samples having MDTs across samples (Fig. 2b). The median fold change in expression between the MDT and the second most abundant transcript of the gene was 5.43.
Comparing the MDTs between each cancer cohort with their normal tissue MDT profiles in GTEx revealed 6114 unique cancer-specific MDTs (cMDTs) from 4151 genes across eight cancer cohorts (Supplementary File 1). 41.0% of cMDTs were sample-specific, occurring only in one sample within a cancer type, while 57.1% of cMDTs were found once within a cohort (Supplementary File 2). Most cMDTs were rare events, observed in less than 10% of samples per cohort (Supplementary Fig. 1). Metastatic epithelial ovarian cancer and metastatic melanoma were found to have the highest number of isoforms-switching events (Fig. 2c). We also observed generally a higher number of MDT switching events in TuPro samples than in TCGA samples (Fig. 2c). Of the 6114 total cMDTs, 230, 301, and 71 cMDTs were commonly found in TCGA and TuPro for ovary, melanoma, and AML respectively (Fig. 2d).
However, 526 cMDTs were found in over 30% of TCGA or TuPro samples (Supplementary File 3). Of these, 32 unique cMDTs were found in TCGA and TuPro cohorts of specific cancer types. Table 1 lists the top 3 most frequent cMDTs for each cancer type.
Table 1 Percentage of top three cMDTs found frequently in TCGA and TuPro samples (NA: Not available).
One of the top 3 cMDTs is an isoform of the cytoskeletal gene Adducin3 (ADD3), which was overexpressed in 82% of ovarian cancer TuPro metastatic samples and 91.3% of TCGA samples. The cMDT has an additional exon, making the protein sequence longer than the normal MDT (Fig. 3a). Inclusion of exon 14 in ADD3 is associated with poor prognosis in lung cancer, likely due to downregulation of the RNA-binding protein QKI-5. Under normal conditions, QKI-5 prevents exon 14 inclusion in ADD3 [46]. Downregulation of QKI-5 can lead to the inclusion of exon 14 and induce cell proliferation in lung cancer.
Fig. 3: Transcript structure (upper panels) and expression (lower panels) of most frequent cMDTs and MDTs across normal and cancer samples.
The GTEx boxplots show the distribution of MDT and cMDT expression across all samples from the corresponding normal sample type. TCGA and TuPro boxplots show cMDT and MDT expression in samples where the transcript is found as cMDT. Structure and expression of (a) ENST00000356080 (cMDT) and ENST00000277900 (MDT) of ADD3, (b) ENST00000614778 (cMDT) and ENST00000264933 (MDT) of PDCD6, (c) ENST00000398571 (cMDT) and ENST00000436269 (MDT) of USP34, and (d) ENST00000335757 (cMDT) and ENST00000407327 (MDT) of SLC442.
In melanoma, we identified the 121-amino-acid (AA) long isoform of programmed cell death 6 (PDCD6), ENST00000614778, as a cMDT, while the longest transcript, ENST00000264933, was an MDT in normal skin samples (Table 1, and Fig. 3b). The normal isoform has 5 EF-hand domains involved in calcium-dependent protein interactions. The cMDT of PDCD6 has only 2 EF-hand domains, which likely disrupts its protein interactions. Incapable of interacting with its binding partners, PDCD6 is likely not able to induce apoptosis by interacting with DAPk1 [47] and inhibit angiogenesis by interacting with VEGFR-2 [48].
In AML, we identified the longest isoform of ubiquitin-specific peptidase 34 (USP34), ENST00000398571, as a cMDT. Overexpression of USP34 has been linked to growth in pancreatic cancer by upregulating phosphorylated AKT and phosphorylated PKC [49] and in laryngeal squamous cell carcinoma by stabilizing SOX2 [50].
Because many diagnostic assays detect transmembrane protein expression, we assessed the percentage of transmembrane genes showing MDT switching events in our cohort. In total, 1707 cMDTs encoding transmembrane proteins were identified, of which 131 had a frequency >30% in at least one cancer cohort. One of these transmembrane proteins was Choline Transporter-Like protein 2 (CTL2), which is involved in cellular choline transport and is encoded by the gene Solute Carrier family 44 member 2 (SLC44A2). We detected frequent MDT switching events for CTL2 in ovarian cancer (Table 1, and Supplementary File 2). In normal ovarian tissue, we detected the 704 AA long CTL2-P1 (ENST00000407327) transcript as MDT, whereas in ovarian cancer, we detected the slightly longer 706 AA long CTL2-P2 (ENST00000335757) as cMDT (Fig. 3d). The two transcripts use different starting exons, resulting in distinct N-terminal regions. The detected cMDT is reported to be more expressed than CTL2-P1 in squamous carcinoma cell lines [51]. Interestingly, CTL2-P2 has detectable choline transport activity, while P1 does not [52]. Since the aberrant choline mechanism is known to drive cancer progression, targeting SLC44A2 isoform switching might be a therapeutic option in ovarian cancer [53, 54]. In addition, 7.88% of TCGA Ovarian samples showed SLC44A2 amplification in cBioPortal (cBioPortal, accessed on 29.07.2024).
Since the spliceosome recognizes pre-mRNA and produces mature mRNA by removing introns and joining exons, we hypothesized that spliceosome dysfunction from mutation or mis-splicing would cause large-scale changes in the overall alternative splicing landscape. To test this hypothesis, we categorised samples by spliceosome splicing and mutational status. Our analysis revealed that samples with a switch, or a switch plus a mutation, in a spliceosome gene consistently had more cMDTs (Supplementary Fig. 2) across melanoma and ovarian cancer samples (p-value < 0.0003). Interestingly, spliceosomes with mutations but no MDT switch showed no significant difference compared with other samples, highlighting the severity of the functional impact of switching events.
Diagnostic cMDTs with peptide evidence in proteomics data
To validate the identification of cMDTs and ensure their expression at the protein level for potential clinical biomarker development, we utilised DIA data from mass spectrometry measurements on matched protein samples. Identifying cMDTs in mass-spectrometry-based proteomics data is challenging because of generally low peptide counts, limited protein sequence coverage, and many shared peptides between isoforms. Hence, we developed a strategy that categorised the proteomics evidence into three levels. First, we focused on cMDT-specific peptides, which are observed only in cMDTs and not in other protein isoforms (Supplementary Files 1 and 4). These cMDT-specific peptides make up Class I cMDTs that have the strongest evidence for the translation of cMDTs. Class I cMDTs are valuable for a simple yet powerful diagnostic assay that probes the expression of a single cMDT isoform for cancer diagnostics. Second, we weakened the exclusiveness criteria and allowed peptides to be present in cMDTs and other transcripts except the normal MDT (Fig. 4). We refer to these cMDTs as Class II. While peptides of Class II cMDTs are not cMDT-specific, they are specific to isoforms that are not dominantly expressed in normal tissue (Supplementary Files 1 and 4). In the last evidence category, we compared the relative expression of multiple peptides of any gene isoform between samples expressing and lacking the cMDT (see ‘Methods’). Although these 3rd-class peptides have the weakest evidence level and do not directly indicate which isoforms are differentially expressed, they still provide evidence that gene splicing is perturbed (Supplementary File 1).
Fig. 4: Proteomics data analysis strategy.
RNA sequencing data from the TuPro projects were used to predict the cMDT in each sample. Class I peptides in pink: cMDT-specific peptides were used to confirm that the cMDT was found in the samples (depicted in pink). Class II peptides: non-MDT peptides found on cMDT and other transcripts but not found on MDT were used to confirm the occurrence of cMDT in the proteomics data (depicted in green). Class III: Differential relative expression of peptides between samples with or without peptide evidence (depicted in black).
At the sample level, we detected a total of 214 ovarian cancers, 271 melanomas, and 27 AML Class I cMDTs (Table 2). In addition to the Class I cMDTs, we detected 355, 479, and 27 Class II cMDTs (Table 2) in ovarian cancer, melanoma, and AML, respectively. Together with Class III peptides, we validated 4028 cMDTs, representing 12.1% of the 33,204 cMDTs identified. Interestingly, ~58–100% of the detected Class I cMDTs and ~22–33% of Class II cMDTs were annotated as principal 1 isoforms in the APRIS database [55] (Supplementary Fig. 3), despite not being expressed as an MDT in GTEx.
Table 2 Summary statistics on all detected cMDTs in the proteomics data with Class I, Class II and Class III evidence (see Supplementary File 1 for detailed list).
To understand why certain Class I and Class II cMDTs were not detected at the proteomics level in specific samples, we compared their RNA expression levels based on whether peptide evidence was detected in the corresponding matched proteomics sample. In ovarian cancer, as shown in Fig. 5, isoforms with peptide evidence display higher RNA expression than transcripts without peptide evidence (Fig. 5, Wilcoxon test p-values: Class I, 1.401e-08; Class II, 8.434e-14). We observed the same result for melanoma and AML samples (Supplementary Figs. 4 and 5) and saw a positive correlation with protein sequence length (Supplementary Fig. 6), suggesting that lowly expressed short cMDT are harder to detect at the protein level.
Fig. 5: Sample-specific RNA expression of Class I and Class II cMDTs according to peptide evidence and their prevalence across matched ovarian cancer samples.
a Density plots of mRNA expression (TPM) of cMDTs with or without cMDT-specific peptide evidence. cMDTs with peptide evidence are highlighted in dark violet, and cMDTs without peptide evidence are depicted in light violet. b Density plots of mRNA expression (TPM) of Class II cMDTs with or without peptide evidence. cMDTs with peptide evidence are represented in dark green, and cMDTs without peptide evidence are shown in light green. P-values were calculated using the Wilcoxon test for two distributions. c The percentage of the top 20 detected Class I cMDTs in the proteomics data. Blue depicts the percentage of cMDTs detected in RNA samples. Dark violet emphasizes the percentage of cMDTs detected with Class I evidence in matched proteomics data. The correlation coefficient between prevalences is 0.63 (p-value: 0.002). d The percentage of the top 20 detected Class I cMDTs in the proteomics data ordered by their RNA expression (TPM) value. Blue shows the percentage of cMDTs detected in RNA samples. The percentage of Class I cMDT is highlighted in dark violet. The correlation coefficient between prevalences is 0.69 (p-value: 0.0007).
Figure 5c shows the prevalence (%) of the top 20 Class I peptides detected in RNA-Seq and matched proteomics data in ovarian cancer. Figure 5d shows the top 20 cMDTs ordered by expression values in the RNA-Seq samples. One of the most frequent Class I cMDTs with an isoform-specific peptide was ADD3-202 (ENST00000356080) (Table 1). We identified ADD3-202 in 79.1% of Ovarian RNA samples and 45% of matched proteomics samples. As ADD3 cMDT is the longest isoform of the gene, we wondered whether we could detect a cMDT-specific peptide in normal samples. The TuPro dataset does not have normal samples like the PaxDB v5 [34]. The PaxDB database contains peptide files from normal proteomics datasets across various tissue types. We searched ovary samples in the PaxDB database (n = 3) and could not identify a cMDT-specific peptide (peptide sequence = LEENHELFSK) in those samples.
Another frequent Class I cMDT peptide belonged to the ATP2B4 transcript ENST00000357681. ATP2B4 encodes a plasma membrane calcium-transporting ATPase 4 enzyme (PMCA4), which releases Ca2+ from cells to help maintain calcium homoeostasis. In pancreatic ductal adenocarcinoma, PMCA4 is overexpressed and promotes apoptotic resistance and cell migration [56]. The longest isoform, ATP2B4-203, was detected in 77.5% of matched-proteomics samples (Fig. 5) and was also observed in 24.4% of TCGA samples.
In melanoma samples, 223 (82.3%) cMDTs with Class I peptides and 432 (90.2%) cMDTs with Class II peptide evidence could also be detected in the TCGA melanoma dataset. In Ovarian samples, 198 (92.5%) cMDTs with Class I peptides and 294 (82.8%) cMDTs with Class II peptides also existed as cMDTs in the TCGA ovarian cancer data. And lastly, 26 of 27 (96.2%) cMDTs with Class I peptides and 25 of 27 (92.6%) cMDTs with Class II peptides were also found in the AML TCGA dataset.
The majority of cMDTs with domain information lose interactions with other proteins and drug compounds
To analyse protein-drug interaction disruptions by cMDTs, we computed isoform-specific protein and drug interaction networks using the STRING [25], 3did [26], PDBsum [35], DrugPort, and Ensembl [36] databases. Using our isoform-specific interaction databases, we observed that cMDTs frequently interact with known cancer-related genes (Supplementary Fig. 7B), as previously shown [27]. 191 or 310 cMDTs are even themselves known as cancer driver genes in COSMIC or OncoKB, respectively.
Among the 1147 genes encoding 1652 cMDTs with domain-domain interaction information, 906 cMDTs lost all interactions with their interaction partners. In contrast, 28 of them lost at least one interaction, and 718 kept all interactions with their protein partners (Fig. 6a, left panel). TRAPPC5 had the most frequent interaction losses in melanoma metastases. The TRAPPC5 protein is an essential subunit of a TRAPP complex, transporting vesicles from the Endoplasmic reticulum to the Golgi [57]. Overexpression of TRAPPC5 is associated with hepatocellular carcinoma progression [58], but the identity of the transcript driving the cancer progression remains unknown. We identified the transcript TRAPPC5-203 (ENST00000595985) in 49.5% of TCGA primary melanoma samples, 60.4% of TCGA metastatic melanoma samples, and 66.7% of TuPro samples. Interestingly, we had identified the same transcript as a frequent interaction-disrupting cMDT in our previous PCAWG study [27, 59]. In contrast to the 188 AA long MDTs in GTEx (TRAPPC5-202: ENST00000426877 or TRAPPC5-201: ENST00000317378), the TRAPPC5-203 cMDT is only 121 AA short, which leads to the loss of all its native interaction partners, including clinically relevant proteins annotated in ClinVar (Fig. 6b). High expression of the cMDT TRAPPC5-203 is associated with lower survival rate in TCGA samples in the GEPIA2 web server. In contrast, high or low expression of GTEx MDT isoforms does not affect survival probability, suggesting the cMDT isoform as an important diagnostic biomarker or drug-target candidate for melanoma (Supplementary Fig. 8).
Fig. 6: Isoform-protein and isoform-drug Interaction databases and example cases for interaction losses.
a Number of cMDT and non-cMDT genes (top), number of cMDT genes with and without domain interaction information in the 3did database (left-middle), cMDTs that lost all interactions, some interactions or keep all interactions compared to STRING canonical isoform (left-bottom), cMDT genes with and without structural information with a drug binding (right-middle), cMDTs that lost or retained interactions with their drugs (right-bottom). b The structure of the TRAPP Complex, and a snapshot of TRAPPC5 isoform interactions. c The structure of the CBP coactivator binding domain and p53 TAD domain, and a snapshot of TP53 interaction. d The cMDT of PDGFRA identified in melanoma samples lacks several extracellular domains and intracellular kinase domains, including the tyrosine kinase domain. The structure of the PDGFRA tyrosine kinase domain in complex with sunitinib (PDB ID: 6JOK) is shown to illustrate the drug-targeted region lost in the cMDT. e The cMDT of EPHB4 identified in TCGA-ovarian samples lacks several extracellular domains and intracellular domains, including the tyrosine kinase domain (PDB ID: 6FNM).
In ovarian cancer, TP53 showed isoform switching in 20 samples (4.8%). The p53 protein is a well-known transcription factor with oncogenic [60] or tumour suppressor roles in most cancer types [61]. Under normal conditions, p53 functions as a central gatekeeper to various cellular functions, including apoptosis and the cell cycle [62,63,64]. In our cohort, 18 samples (1 AML, 17 Ovarian samples) expressed the 261 AA long TP53-209 (ENST00000504937) cMDT. In normal GTEx samples, the predominant MDT was TP53-201 (ENST00000269305). TP53-209 encodes the Δ133p53α isoform, which is produced from an alternative promoter region of TP53. The resulting protein product lacks two transactivation domains (TAD1 and TAD2), a proline-rich domain (PRD), and parts of a DNA-binding domain [65, 66]. According to our isoform-specific protein interaction database, Δ133p53α cannot interact with critical protein partners, including MDM2, MDM4, POLR2E, and EP300 (Fig. 6c), suggesting its functional impact at the cellular level.
Interestingly, the cMDT we discovered were located near known cancer-related genes. We also found that 191 cMDTs were COSMIC cancer census genes themselves. Compared with cancer genes in OncoKB, this number increases to 310, confirming our earlier observations about the proximity of cMDTs to cancer-associated genes (Supplementary Fig. 7B).
Alternative splicing cannot only disrupt protein-protein interactions but also induce resistance to drug therapies. 641 cMDTs across our cancer cohort had drug interaction information according to our new IsoDrug database (Fig. 6a, right panel). The canonical protein isoforms of the 641 cMDTs are known targets of 166 drugs, resulting in 1070 isoform-drug interactions. 56% of cMDTs lose their binding capability to their potential drug targets (Fig. 6a, right panel). In those cases, the patient likely resists the drug, showing no response because the target protein lacks the binding site.
Because many cancer-related drugs, such as Imatinib and Sunitinib, are tyrosine-kinase inhibitors (TKIs), we analysed whether any tyrosine kinase in our cohort expresses a cMDT that lacks the binding site for a TKI compound. In total, we found 15 cMDTs whose canonical isoform binds to one of 11 TKI compounds. The most frequent cMDT was PDGFRA-210 (ENST00000512522), a 154 AA-long isoform of PDGFRA expressed predominantly in TuPro melanoma samples. The PDGFRA cMDT lacks tyrosine kinase domains and four Ig-like domains, where the former domain is the primary target of Sunitinib (PDB ID: 6JOK) (Fig. 6d). Thus, it is likely that melanoma patients expressing PDGFRA-210 as the predominant isoform might not respond to Sunitinib treatment even if the protein is mutated. Another similar example is the EPHB4-210 cMDT (ENST00000616502), which was found in 26 TCGA ovarian cancer samples. The tyrosine kinase receptor EPHB4 is overexpressed in ovarian cancer and other cancer types and is activated by Ephrin-B2 [67]. The binding of Ephrin-B2 to EPHB4 activates downstream targets that promote angiogenesis. Interestingly, the short cMDT EPHB4-210 has only the ligand-binding domain and lacks all other domains, including the entire tyrosine kinase domain (Fig. 6e). In contrast, associated GTEx samples show evidence of a longer isoform EPHB4-201 (ENST00000358173) in normal samples.

