why choose us

300×250 Ad Slot

Research Article: Characterization and clustering of immunocompetent proteins based on ESM-2 sequence embedding and physicochemical feature fusion

Date Published: 2026-09-30

Abstract:
Traditional sequence-comparison methods face challenges when classifying and identifying low-homology proteins formed by convergent evolution. In this exploratory study, we built a PCA-based unsupervised clustering workflow combining high-dimensional embeddings from the frozen pre-trained protein-language model ESM-2 and weighted-fused physicochemical features. The workflow was explored for family-level grouping, identification of putative functional conservation, and cross-species evolutionary analysis of immune-active proteins. An independent external test dataset was further used to preliminarily probe its potential generalization behaviour. Nineteen arthropod-derived immunocompetent proteins (tick lipocalin antihistamine proteins, mosquito D7 salivary proteins, dust mite allergens, and two additional arthropod immune-active proteins) formed our reference dataset. Eleven test sequences, comprising nine independent vertebrate Cathelicidin peptides and two tick lipocalin positive controls, including nine vertebrate cathelicidin antimicrobial peptides and two tick lipocalin homologs, served as the external test dataset. Using Python3.9, we implemented an analytical pipeline. The ESM-2 t33_650M_UR50D model in frozen pre-trained mode generated 1280-dimensional sequence embeddings. Six core physicochemical and secondary-structure descriptors were computed. An empirically-determined weight ratio (0.7:0.3) was applied to build a weighted fused dissimilarity matrix. After full feature concatenation, PCA-based clustering and visualization were performed. On the reference data set, the fusion features showed a significant visual separation trend among different functional protein families. The fusion dissimilarity value within the family is usually lower than the dissimilarity value between the families, and there is a statistically significant difference at the group level (P < 0.05). The 9 vertebrate cathelicidin sequences in the test dataset form a unique cluster separated from the arthropod reference protein. The observed clustering patterns are generally consistent with the evolutionary relationship of species and the calculated physical and chemical properties. This exploratory ESM-2-embedding and physicochemical-feature-fusion PCA clustering workflow shows promising descriptive performance for immune-active proteins. It can generate family-level grouping without relying on high-quality multiple-sequence alignments and may help analyse low-homology protein datasets. Nevertheless, given dataset-size limitations and the absence of quantitative clustering-validation metrics, its broad-scale classification accuracy, stability and generalizability remain to be further validated on larger datasets. This pipeline provides a reproducible bioinformatic framework for exploratory family screening, evolutionary interpretation and functional hypothesis generation for animal-derived immune-active proteins.

Introduction:
Protein serve as the core functional carriers of life activities, and its functional conservation represents a key molecular basis for maintaining adaptive evolution and physiological homeostasis in species ( 1 – 3 ). Over the course of long-term natural selection and species divergence, homologous functional proteins across different species—and even among various developmental stages of the same species—typically exhibit evolutionary characteristics indicative of sequence heterogeneous differentiation and…

Read more

300×250 Ad Slot