{
    "componentChunkName": "component---src-templates-article-page-js",
    "path": "/journals/biology/micropub-biology-002204",
    "result": {"data":{"article":{"manuscript":{"id":"72d47e77-f90e-4828-9273-b88b9ce9969f","submissionTypes":["new finding"],"citations":[],"doi":"10.17912/micropub.biology.002204","dbReferenceId":null,"pmcId":null,"pmId":null,"proteopedia":null,"reviewPanel":null,"species":["human"],"integrations":[],"corrections":null,"history":{"received":"2026-05-15T18:17:09.089Z","revisionReceived":"2026-08-14T00:18:45.204Z","accepted":"2026-08-31T22:26:20.457Z","published":"2026-09-04T20:59:57.368Z","indexed":"2026-09-18T20:59:57.368Z"},"versions":[{"id":"dc83d15d-9759-4e42-a291-de7c340161e3","decision":"revise","abstract":"<p>Creutzfeldt-Jakob Disease (CJD) is a 100% fatal prion disorder with no approved treatments. Effective therapeutics must simultaneously cross the Blood-Brain Barrier, selectively target pathogenic PrPSc, and neutralize prion propagation. This study developed a machine learning-driven pipeline trained on 21 anti-prion antibodies to rank 25,000+ sequences and generate novel variants. The models demonstrated strong performance (Neutralization AUC = 0.92; Specificity AUC = 0.78), identified 10 high-percentile candidates, and confirmed prion-specific learning using FDA CNS antibodies as external controls.</p>","acknowledgements":"<p>The authors would like to thank Academy of the Canyons for providing resources to enable this project.</p>","authors":[{"affiliations":["College of the Canyons, Santa Clarita, California, United States","Academy of the Canyons, Santa Clarita, California, United States"],"departments":["",""],"credit":["conceptualization","dataCuration","formalAnalysis","fundingAcquisition","investigation","methodology","project","resources","software","validation","visualization","writing_originalDraft","writing_reviewEditing"],"email":"vtiruvellore@my.canyons.edu","firstName":"Vinayan","lastName":"Tiruvellore","submittingAuthor":true,"correspondingAuthor":true,"equalContribution":false,"WBId":null,"orcid":"https://orcid.org/0009-0003-5451-2503"},{"affiliations":["Academy of the Canyons, Santa Clarita, California, United States","College of the Canyons, Santa Clarita, California, United States"],"departments":["",""],"credit":["supervision","resources"],"email":"mkoegle@hartdistrict.org","firstName":"Mike","lastName":"Koegle","submittingAuthor":false,"correspondingAuthor":false,"equalContribution":false,"WBId":null,"orcid":null}],"awards":[],"conflictsOfInterest":"<p>The authors declare that there are no conflicts of interest present.</p>","dataTable":{"url":null},"extendedData":[],"funding":"<p>This work did not receive external funding. Support was provided by the authors and Academy of the Canyons, Santa Clarita, CA.</p>","image":{"url":"https://portal.micropublication.org/uploads/704d795891ef24bc74c9e5699dd42d7f.png"},"imageCaption":"<p>A) Machine learning pipeline for anti-CJD antibody candidate discovery. Schematic overview of the computational workflow used to evaluate antibody sequences against three therapeutic criteria: blood–brain barrier penetration, PrPSc specificity, and predicted neutralization activity. Background antibody sequences were scored relative to curated anti-prion exemplars, ranked by composite therapeutic potential, and used as the basis for mutation-based generation of novel antibody variants.</p><p>B) Feature engineering and sequence-level model inputs. Quantitative physicochemical descriptors were extracted from antibody CDRH3 sequences, including length, hydropathy, residue composition, charge distribution, aromatic fraction, and glycine/proline content. These engineered sequence features were used to train classifiers that distinguish anti-prion antibodies from non-prion background sequences and support multi-objective ranking.</p><p>C) Model performance and candidate ranking results. Predicted neutralization and selectivity scores separated curated anti-prion positive exemplars from background antibodies and FDA-approved CNS antibody controls. ROC AUC and Mann–Whitney U test results show statistically significant discrimination between anti-prion and non-prion sequences, while ranked candidate pool analysis highlights enrichment of high-scoring antibody candidates among the top-ranked sequences.</p><p>D) Structural comparison of parent and mutation-derived antibody candidates. ESMFold-predicted CDRH3 structures show conformational effects of single-residue deletions in high-scoring mutation-derived variants. Removal of structurally restrictive or bulky residues, including proline and tyrosine, altered loop geometry and corresponded with improved composite therapeutic ranking relative to the original parent sequences.</p>","imageTitle":"<p><i>Fig. 1 (A, B, C, D)</i></p>","methods":"<p></p>","reagents":"<p></p>","patternDescription":"<p>CJD is a rare but universally fatal prion disease. There is no cure, no disease-modifying therapy. CJD is caused by the misfolding of native prion protein, PrPC, into the pathogenic isoform PrPSc. PrPSc acts as a self-propagating template, converts normal proteins into misfolded forms, forms amyloid aggregates resistant to degradation, and causes widespread neuronal death. Prion diseases represent one of the most challenging classes of neurodegenerative disorders because the infectious agent is a misfolded protein and not a virus or bacterium. PrPSc is structurally similar to native PrPC, the central nervous system is protected by the highly selective BBB, therapeutics must avoid disrupting normal prion protein function, and despite decades of research, no clinically approved antibody or small-molecule therapy exists.</p><p>Machine learning approaches to antibody engineering have shown promise in identifying candidate therapeutics. By training ranking models on the physicochemical features of known anti-prion antibodies, it is possible to evaluate large libraries of candidates and identify those most likely to satisfy therapeutic criteria. The experimental question was: Can a machine learning-driven computational pipeline propose antibody candidates optimized for anti-CJD prion proteins that satisfy all three therapeutic criteria: BBB penetration, specificity, and neutralization? The hypothesis was that a machine learning pipeline trained on physicochemical features of known anti-prion antibodies would rank true anti-prion candidates significantly above background sequences, distinguish prion-specific antibodies from unrelated CNS antibodies, and generate novel, high-scoring variants through mutation modeling.</p><p>The primary control was baseline background antibodies. An external negative control was also included using FDA-approved CNS-targeting antibodies. These were used to ensure the model captured prion-specific sequence signatures rather than general CNS-targeting characteristics. For data collection, 21 curated exemplar anti-prion antibodies were used as positives, and more than 20,000 non-prion sequences were used as background. For feature engineering, CDRH3 sequences were converted into quantitative descriptors, including 3-mer motif embeddings, amino acid composition, length, net charge at pH 7.4, hydropathy using the Kyte-Doolittle scale, aromatic fraction, and residue class proportions.</p><p>Model training was designed to distinguish positives from background. The pipeline included a neutralization classifier, a specificity classifier, and a BBB proxy scoring model. Candidate ranking was then performed by scoring more than 20,000 sequences and calculating a composite score across all three criteria. The top 25 candidates were selected. Mutation modeling was then applied through targeted sequence modifications, and the modified sequences were re-scored to identify novel high-ranking candidates.</p><p>The neutralization classifier achieved ROC AUC = 0.9239 ± 0.0810, and the specificity classifier achieved ROC AUC = 0.7759 ± 0.1217. Both were substantially above random expectation, confirming that CDRH3 physicochemical features carry meaningful discriminative signal for anti-prion activity. The BBB penetration model was a literature-based proxy. Strong performance was achieved despite training on only 21 curated exemplars.</p><p>Mann-Whitney U tests confirmed statistically significant score separation between positive exemplars and background sequences and between positives and FDA CNS antibodies. Positives versus background for neutralization had a U statistic of 180000 and a p-value of 0.000000102. Positives versus background for specificity had a U statistic of 300000 and a p-value of 0.0000000001. Positives versus FDA CNS antibodies for neutralization had a U statistic of 72 and a p-value of 0.0003129. Positives versus FDA CNS antibodies for specificity had a U statistic of 120 and a p-value of 0.0000609. These extremely low p-values indicate the differences were highly unlikely to occur by chance and provided rigorous statistical validation of the machine learning predictions.</p><p>The neutralization score box plot for the top 50 candidates compared predicted neutralization scores for positive exemplars and background sequences. Anti-prion antibodies clustered near the highest predicted neutralization scores, around 0.99 to 1.0. Background sequences remained spread across lower score values, around 0.3 to 0.7. This demonstrated strong model separation between prion-targeting and non-prion sequences. The selectivity score distribution showed predicted selectivity scores for distinguishing pathogenic PrPSc targeting. Positive exemplar antibodies again occupied the highest scoring region, around 0.99 to 1.0, while background sequences occupied lower score ranges, around 0.07 to 0.8. This confirmed that the classifier identified prion-specific sequence features.</p><p>Neutralization scores were also compared across known anti-prion antibodies, FDA CNS antibodies, and background sequences. Positive exemplars clustered near maximal scores, around 1.0, indicating strong predicted neutralization potential. FDA CNS antibodies scored near zero, confirming the model did not simply favor antibodies used in CNS therapies. Background sequences again showed a low distribution of scores. Selectivity scores showed the same pattern. Anti-prion exemplar antibodies clustered near the highest scoring region, while FDA CNS antibodies and most background sequences remained near zero, indicating minimal predicted prion specificity. The separation between groups supports that the classifier captures prion-specific sequence signatures rather than general antibody features.</p><p>The composition of the ranked candidate pool showed the size of the sequence libraries evaluated by the ranking pipeline. The candidate pool consisted of more than 20,000 original antibody sequences and about 7,000 mutation-generated variants. This large search space allowed the machine learning models to identify rare high-scoring candidates within diverse antibody sequence space. Ten candidates ranked in the 95th to 99th percentile across 29,574 evaluated sequences. Positive exemplars were heavily enriched in the top 50, while FDA CNS antibodies were entirely absent. Mutation modeling generated 7,309 novel sequences not present in the original library. The top novel candidates ranked 3rd, 7th, and 12th overall, outperforming 99.97% of all evaluated sequences.</p><p>Single-residue deletions produced the largest improvements. In one case, parent sequence QQNKNWPPGT ranked 17, while novel sequence QQNKNWPGT ranked 3. A single Pro8 deletion distinguished the parent from the novel generated sequence. Proline is structurally unique among amino acids because its cyclic side chain rigidly constrains the peptide backbone, forcing a kink in the loop conformation. Removing Pro8 relieved this constraint, allowing the CDRH3 loop to adopt a more flexible, extended conformation visible in the ESMFold prediction. This conformational shift improved the composite therapeutic score from 0.9738 to 0.9810, advancing the sequence from rank 17 to rank 3 out of 29,574 total candidates. This candidate was the highest ranked mutant or novel variant out of around 7,000.</p><p>In another case, parent sequence SSYTITNTQK ranked 367, while novel sequence SSTITNTQK ranked 12. A single Tyr3 deletion distinguished the parent from the novel generated sequence. Tyrosine carries a bulky aromatic ring that projects outward from the peptide backbone, imposing steric constraints on the surrounding loop geometry, which was clearly visible in the ESMFold structure of the parent. Its removal produced a markedly flatter, more compact loop architecture, reducing steric bulk in a region predicted to be relevant for PrPSc engagement. This structural change drove the largest score improvement across all three mutation pairs. The combined score rose from 0.9439 to 0.9759, a +0.0320 improvement, and rank jumped from 367 to 12, which was a 355-position improvement from a single residue deletion.</p><p>This study demonstrates that machine learning can effectively identify and optimize antibody therapeutic candidates for Creutzfeldt-Jakob Disease despite extremely limited training data. Both predictive classifiers successfully distinguished anti-prion antibodies from unrelated background sequences using only sequence-derived physicochemical descriptors. The neutralization model achieved strong discriminative performance, while the specificity classifier demonstrated moderate yet statistically significant predictive capability. These results indicate that meaningful therapeutic signal is encoded within antibody sequence features such as residue composition, charge distribution, and hydrophobicity.</p><p>Composite ranking across three independent therapeutic constraints, including blood-brain barrier penetration, prion-specific binding, and neutralization potential, allowed the pipeline to approximate real-world drug development trade-offs. Rather than optimizing a single metric, the framework balances multiple biological requirements simultaneously, reflecting the complexity of CNS therapeutic development. This multi-objective scoring approach enabled identification of candidates that perform well across all criteria rather than excelling in only one dimension.</p><p>External validation using FDA-approved CNS antibodies provided additional evidence that the models capture prion-specific sequence patterns rather than generic properties of antibodies capable of entering the brain. These control antibodies scored near zero in both predictive models, while anti-prion exemplars clustered near the highest score ranges. This separation demonstrates that the classifiers learned biologically relevant signal associated with prion targeting rather than simply recognizing general CNS therapeutic traits.</p><p>The mutation model further expanded the search space by generating targeted sequence modifications around top-ranking candidates. Guided amino acid substitutions preserved biochemical compatibility while introducing structural diversity, enabling exploration of nearby sequence space. Several mutation-derived variants achieved scores comparable to or exceeding their parent sequences, demonstrating the potential for computational optimization of antibody therapeutics.</p><p>Traditional anti-prion drug discovery is bottlenecked by the scarcity of validated antibody datasets and the difficulty of experimental prion screening. This pipeline demonstrates that sequence-level machine learning can extract meaningful biological signal from as few as 21 positive examples. The multi-constraint scoring framework is disease-agnostic, and the same architecture could be retrained for Alzheimer’s, Parkinson’s, and other protein misfolding disorders where BBB penetration and isoform selectivity are simultaneously required. The identified candidates represent a prioritized set for downstream validation, including molecular docking against PrPSc crystal structures, binding affinity assays, cell-based BBB transcytosis models, and ultimately prion cell model neutralization testing. Future work would center around gathering more positive exemplars, wet-lab validation, expansion of the mutation model, and replacing the BBB proxy with a trained model so that the third scoring dimension becomes more rigorous.</p>","references":[{"reference":"Aguzzi A, Calella AM. 2009. Prions: Protein aggregation and infectious diseases. Physiological Reviews. 89: 1105.","pubmedId":"","doi":"10.1152/physrev.00006.2009"},{"reference":"Dobson CL, Devine PW, Phillips JJ, Higazi DR, Lloyd C, Popplewell AG. 2016. Engineering the surface properties of a human monoclonal antibody prevents self-association and rapid clearance in vivo. Scientific Reports. 6: 38644.","pubmedId":"","doi":"10.1038/srep38644"},{"reference":"Mieczkowski C, Zhang X, Lee D, Nguyen K, Lv W, Wang Y, et al., Gries JM. 2023. Blueprint for antibody biologics developability. mAbs. 15: 2185924.","pubmedId":"","doi":"10.1080/19420862.2023.2185924"},{"reference":"Pankiewicz JE, Sanchez S, Kirshenbaum K, Kascsak RB, Kascsak RJ, Sadowski MJ. 2019. Anti-prion protein antibody 6D11 restores cellular proteostasis of prion protein through disrupting recycling propagation of PrPSc and targeting PrPSc for lysosomal degradation. Molecular Neurobiology. 56: 2073.","pubmedId":"","doi":"10.1007/s12035-018-1208-4"},{"reference":"Prusiner SB. 1998. Prions. Proceedings of the National Academy of Sciences. 95: 13363.","pubmedId":"","doi":"10.1073/pnas.95.23.13363"},{"reference":"Ruiz Lopez E. 2021. Transportation of single-domain antibodies through the blood-brain barrier. Biomolecules. 11: 1131.","pubmedId":"","doi":"10.3390/biom11081131"},{"reference":"Triguero D, Buciak JB, Pardridge WM. 1989. Blood-brain barrier transport of cationized immunoglobulin G. Proceedings of the National Academy of Sciences. 86: 4761.","pubmedId":"","doi":"10.1073/pnas.86.12.4761"},{"reference":"Uger MD, Chai V, Ciolfi V, Cashman NR, Tian B, Wong WY, Chao HLM. 2016. Antibodies and conjugates that target misfolded prion protein.","pubmedId":"","doi":""},{"reference":"Williamson RA, Burton DR, Prusiner SB. 2001. Antibodies specific for native PrPSc.","pubmedId":"","doi":""},{"reference":"Yang W. 2014. Efficient penetration of the blood-brain barrier by anti-prion antibody fragments. Science Translational Medicine. 6: 234ra60.","pubmedId":"","doi":"10.1126/scitranslmed.3008870"},{"reference":"Collinge J, Hawke S. 2009. Prion inhibition.","pubmedId":"","doi":""}],"title":"<p>Against All Odds: Machine Learning Ranking and Generation of Antibody Therapeutics for Creutzfeldt-Jakob Disease</p>","reviews":[],"curatorReviews":[]},{"id":"37ac4ad9-9764-4340-97e5-28fcc115c404","decision":"revise","abstract":"<p>Creutzfeldt-Jakob Disease (CJD) is a 100% fatal prion disorder with no approved treatments. Effective anti-prion candidates must simultaneously cross the Blood-Brain Barrier, selectively target pathogenic PrPSc, and neutralize prion propagation. This study developed a machine learning-driven pipeline trained on 21 anti-prion antibodies to rank 29,574 original and mutation-generated candidate antibody sequences using neutralization, specificity, and a literature-based BBB proxy. The models demonstrated internal performance (Neutralization AUC = 0.9239; Specificity AUC = 0.7759), identified 10 high-percentile computational candidates, generated 7,309 novel variants, and supported prion-specific learning using FDA CNS antibodies as external controls for planned future experimental validation of anti-prion activity.</p>","acknowledgements":"<p>The authors would like to thank Academy of the Canyons for providing resources to enable this project.</p>","authors":[{"affiliations":["College of the Canyons, Santa Clarita, California, United States","Academy of the Canyons, Santa Clarita, California, United States"],"departments":["",""],"credit":["conceptualization","dataCuration","formalAnalysis","fundingAcquisition","investigation","methodology","project","resources","software","validation","visualization","writing_originalDraft","writing_reviewEditing"],"email":"vtiruvellore@my.canyons.edu","firstName":"Vinayan","lastName":"Tiruvellore","submittingAuthor":true,"correspondingAuthor":true,"equalContribution":false,"WBId":null,"orcid":"https://orcid.org/0009-0003-5451-2503"},{"affiliations":["Academy of the Canyons, Santa Clarita, California, United States","College of the Canyons, Santa Clarita, California, United States"],"departments":["",""],"credit":["supervision","resources"],"email":"mkoegle@hartdistrict.org","firstName":"Mike","lastName":"Koegle","submittingAuthor":false,"correspondingAuthor":false,"equalContribution":false,"WBId":null,"orcid":null}],"awards":[],"conflictsOfInterest":"<p>The authors declare that there are no conflicts of interest present.</p>","dataTable":{"url":null},"extendedData":[],"funding":"<p>This work did not receive external funding. Support was provided by the authors and Academy of the Canyons, Santa Clarita, CA.</p>","image":{"url":"https://portal.micropublication.org/uploads/704d795891ef24bc74c9e5699dd42d7f.png"},"imageCaption":"<p>A) Machine learning pipeline for anti-CJD antibody candidate discovery. Schematic overview of the computational workflow used to evaluate antibody sequences against three screening criteria: blood–brain barrier penetration, PrPSc specificity, and predicted neutralization activity. Background antibody sequences were scored relative to curated anti-prion exemplars, ranked by composite candidate-prioritization score, and used as the basis for mutation-based generation of novel antibody variants.</p><p>B) Feature engineering and sequence-level model inputs. Quantitative physicochemical descriptors were extracted from antibody CDRH3 sequences, including length, hydropathy, residue composition, charge distribution, aromatic fraction, and glycine/proline content. These engineered sequence features were used to train classifiers that distinguish anti-prion antibodies from non-prion background sequences and support multi-objective ranking.</p><p>C) Model performance and candidate ranking results. Predicted neutralization and selectivity scores separated curated anti-prion positive exemplars from background antibodies and FDA-approved CNS antibody controls. ROC AUC and Mann–Whitney U test results show statistically significant discrimination between anti-prion and non-prion sequences, while ranked candidate pool analysis highlights enrichment of high-scoring antibody candidates among the top-ranked sequences.</p><p>D) Structural comparison of parent and mutation-derived antibody candidates. ESMFold-predicted CDRH3 structures show conformational effects of single-residue deletions in high-scoring mutation-derived variants. Removal of structurally restrictive or bulky residues, including proline and tyrosine, altered loop geometry and corresponded with improved composite computational ranking relative to the original parent sequences.</p>","imageTitle":"<p><i>Fig. 1 (A, B, C, D)</i></p>","methods":"<p><b><i>Data collection:</i></b> This study was designed as a preliminary computational screening pipeline for candidate anti-prion antibody sequences rather than as evidence of therapeutic efficacy. The positive set consisted of 21 curated anti-prion CDRH3 sequences compiled from patents and literature-derived criteria sets: 8 sequences in the neutralization set, 11 in the selectivity set, and 2 BBB-related comparator sequences. Because experimentally validated anti-prion antibody datasets are scarce, this curated set was used as a small hypothesis-generating positive reference set. For model contrast, 16,163 non-prion human IGH background CDRH3 sequences were used as negatives during training.</p><p><b><i>Control construction:</i></b> The primary negative control was the non-prion IGH background set. An external comparator set consisting of 8 FDA-approved CNS-related antibodies that were not designed to bind prions was used during downstream evaluation. These FDA CNS antibodies were treated as non-prion comparators to test whether the scoring framework was capturing prion-related sequence patterns rather than general CNS-targeting characteristics.</p><p><b><i>Feature engineering:</i></b> Each CDRH3 sequence was converted into a hybrid feature representation consisting of character-level 3-mer sequence features and 7 physicochemical features. The physicochemical features were length, hydrophobic fraction, aromatic fraction, positive-residue fraction, negative-residue fraction, net charge per residue, and glycine/proline fraction. In the final target feature matrix, the 21 positive sequences were represented by 137 distinct 3-mer features and 7 physicochemical features. This representation was chosen to preserve short sequence motif information while also capturing coarse biochemical properties relevant to antibody recognition and developability.</p><p>Machine-learning approach: Two binary classifiers were trained: a neutralization-like classifier and a selectivity-like classifier. Both models used L2-regularized logistic regression with class balancing and the liblinear solver. Positive labels were assigned from the curated anti-prion sequence set, and negative labels were assigned from the non-prion background set. The models therefore estimated whether a sequence more closely resembled the curated neutralization or selectivity exemplars than the background repertoire.</p><p>Training and validation: Model performance was evaluated using 5-fold stratified cross-validation with ROC AUC and average precision. The neutralization-like classifier achieved mean ROC AUC = 0.9239 ± 0.0810, and the selectivity-like classifier achieved mean ROC AUC = 0.7759 ± 0.1217 across folds. Because only 21 curated positive exemplars were available, these models should be interpreted as preliminary ranking models rather than definitive predictive classifiers. The larger variance observed for the selectivity model is consistent with reduced stability under limited positive-sample conditions.</p><p><b><i>BBB proxy scoring model:</i></b> The BBB proxy scoring model was a hand-engineered sequence-ranking function rather than an experimentally trained BBB penetration model. For each antibody sequence, the model computed six physicochemical descriptors: sequence length, approximate molecular weight, net charge at pH 7.4, estimated isoelectric point (pI), GRAVY hydropathy (Kyte-Doolittle average hydropathy), and a hydrophobic patch score defined as the fraction of 6-residue windows containing at least 4 hydrophobic residues. The proxy then assigned a BBB compatibility score by rewarding sequences with net charge near 6.0 and pI near 9.6 using Gaussian-like soft-band terms, while penalizing excessive charge, excessively high pI, large molecular weight, high hydropathy, and high hydrophobic patchiness using smooth soft-penalty terms. The raw score was computed as:</p><p><br><i>0.35 x charge_reward + 0.25 x pI_reward - 0.15 x size_penalty - 0.15 x gravy_penalty - 0.10 x patch_penalty - 0.10 x charge_extreme_penalty - 0.10 x pI_extreme_penalty</i></p><p><br>and then transformed to a 0-1 scale with the logistic function where <i>BBB_score = 1 / (1 + e<sup>(-4·(raw - 0.3)</sup>))</i>. Higher scores indicated stronger predicted BBB compatibility under this proxy, whereas lower scores indicated weaker compatibility. Because no independent BBB transcytosis or transport training dataset was used, this score should be interpreted strictly as a computational ranking proxy rather than direct evidence of BBB penetration.</p><p><b><i>Statistical adequacy and overfitting</i></b>: The positive reference set was small, so overfitting risk is an important limitation. To mitigate this, the classifiers used regularized logistic regression, class balancing, and 5-fold stratified cross-validation, and were evaluated against both a large background repertoire and an external FDA CNS comparator set. These steps support the use of the framework for preliminary prioritization, but they do not eliminate the need for independent experimental validation. Accordingly, all outputs should be interpreted as hypothesis-generating candidate rankings rather than validated therapeutic predictions.</p><p><b><i>Candidate ranking</i></b>: Candidate sequences were ranked by combining the neutralization-like score, selectivity-like score, and BBB proxy score. Ranking was performed on original seed sequences and mutation-generated variants. Across 29,574 evaluated sequences in the ranking/validation workflow, the top candidates were prioritized based on combined percentile performance across the three criteria.</p><p><b><i>Statistical testing:</i></b> Mann-Whitney U tests were used to compare score distributions between predefined groups, including positives versus background and positives versus FDA CNS comparators, for both neutralization-like and selectivity-like scores. These tests were used as score-separation analyses rather than proof of causal biological discrimination.</p><p><b><i>Mutation modeling:</i></b> Sequence diversification was performed computationally by targeted mutation and rescoring. This procedure generated 7,309 novel sequences not present in the original library. These mutated candidates were evaluated using the same ranking framework as the original sequences.</p><p><b><i>Data and code availability</i></b>: The curated input datasets, extracted feature tables, trained scoring artifacts, ranked candidate outputs, and analysis code are publicly available at https://github.com/Axy-lis/CJD_ML_AAO</p>","reagents":"<p></p>","patternDescription":"<p>CJD is a rare but universally fatal prion disease. There is no cure, no disease-modifying therapy. CJD is caused by the misfolding of native prion protein, PrPC, into the pathogenic isoform PrPSc. PrPSc acts as a self-propagating template, converts normal proteins into misfolded forms, forms amyloid aggregates resistant to degradation, and causes widespread neuronal death (Prusiner 1998; Aguzzi and Calella 2009). Prion diseases represent one of the most challenging classes of neurodegenerative disorders because the infectious agent is a misfolded protein and not a virus or bacterium. PrPSc is structurally similar to native PrPC, the central nervous system is protected by the highly selective BBB, candidate interventions must avoid disrupting normal prion protein function, and despite decades of research, no clinically approved antibody or small-molecule therapy exists (Aguzzi and Calella 2009; Pankiewicz et al. 2019).</p><p>Machine learning approaches to antibody engineering have shown promise in identifying candidate antibodies. By training ranking models on sequence-derived CDRH3 features from curated anti-prion antibody exemplars, it is possible to evaluate large libraries of candidates and identify those most likely to satisfy computational screening criteria. The experimental question was: Can a machine learning-driven computational pipeline prioritize candidate anti-prion antibody sequences for CJD that satisfy all three screening criteria: BBB penetration, specificity, and neutralization? The hypothesis was that a machine learning pipeline trained on sequence-derived features from known anti-prion antibodies would rank true anti-prion candidates significantly above background sequences, distinguish prion-specific antibodies from unrelated CNS antibodies, and generate novel, high-scoring variants through mutation modeling.</p><p>The primary control was baseline background antibodies. An external negative control used FDA-approved CNS-targeting antibodies to test whether the model captured prion-specific sequence signatures rather than general CNS-targeting characteristics. For data collection, 21 curated exemplar anti-prion antibodies were used as positives, and more than 20,000 non-prion sequences were used as background. Anti-prion antibody examples and prion-targeting antibody concepts were drawn from prior literature and patent sources (Pankiewicz et al. 2019; Uger et al. 2016; Williamson et al. 2001; Collinge and Hawke 2009). For feature engineering, CDRH3 sequences were converted into quantitative descriptors, including 3-mer motif embeddings, amino acid composition, length, net charge at pH 7.4, hydropathy using the Kyte-Doolittle scale, aromatic fraction, and residue class proportions.</p><p>Model training was designed to distinguish positives from background. The pipeline included a neutralization classifier, a specificity classifier, and a BBB proxy scoring model. Candidate ranking was performed by scoring more than 20,000 sequences and calculating a composite score across all three criteria. The top 25 candidates were selected. Mutation modeling was applied through targeted sequence modifications, and modified sequences were re-scored to identify novel high-ranking candidates.</p><p>The neutralization classifier achieved ROC AUC = 0.9239 ± 0.0810, and the specificity classifier achieved ROC AUC = 0.7759 ± 0.1217. Both were substantially above random expectation, supporting that CDRH3 sequence-derived features carry meaningful discriminative signal for anti-prion activity. The BBB penetration model was a literature-based proxy informed by prior work on antibody BBB transport and anti-prion antibody fragments (Triguero et al. 1989; Yang et al. 2014; Ruiz-López et al. 2021). Strong internal model performance was achieved despite training on only 21 curated exemplars.</p><p>Mann-Whitney U tests showed statistically significant score separation between positive exemplars and background sequences and between positives and FDA CNS antibodies. Positives versus background for neutralization had a U statistic of 180000 and a p-value of 0.000000102. Positives versus background for specificity had a U statistic of 300000 and a p-value of 0.0000000001. Positives versus FDA CNS antibodies for neutralization had a U statistic of 72 and a p-value of 0.0003129. Positives versus FDA CNS antibodies for specificity had a U statistic of 120 and a p-value of 0.0000609. These low p-values support score separation between predefined groups, but do not independently establish antibody activity or guarantee model generalizability.</p><p>The neutralization score box plot for the top 50 candidates compared predicted neutralization scores for positive exemplars and background sequences. Anti-prion antibodies clustered near the highest predicted neutralization scores, around 0.99 to 1.0. Background sequences remained spread across lower score values, around 0.3 to 0.7. The selectivity score distribution showed predicted selectivity scores for distinguishing pathogenic PrPSc targeting. Positive exemplar antibodies again occupied the highest scoring region, around 0.99 to 1.0, while background sequences occupied lower score ranges, around 0.07 to 0.8. This supported model separation between prion-targeting and non-prion sequences.</p><p>Neutralization scores were also compared across known anti-prion antibodies, FDA CNS antibodies, and background sequences. Positive exemplars clustered near maximal scores, FDA CNS antibodies scored near zero, and background sequences showed a low score distribution. Selectivity scores showed the same pattern. The separation between groups supports that the classifier captures prion-specific sequence signatures rather than general antibody features.</p><p>The ranked candidate pool consisted of more than 20,000 original antibody sequences and about 7,000 mutation-generated variants. This search space allowed the machine learning models to identify rare high-scoring candidates within diverse antibody sequence space. Ten candidates ranked in the 95th to 99th percentile across 29,574 evaluated sequences. Positive exemplars were heavily enriched in the top 50, while FDA CNS antibodies were entirely absent. Mutation modeling generated 7,309 novel sequences not present in the original library. The top novel candidates ranked 3rd, 7th, and 12th overall, outperforming 99.97% of all evaluated sequences in this computational screen.</p><p>Single-residue deletions produced the largest improvements. In one case, parent sequence QQNKNWPPGT ranked 17, while novel sequence QQNKNWPGT ranked 3. A single Pro8 deletion distinguished the parent from the novel generated sequence. Proline is structurally unique among amino acids because its cyclic side chain rigidly constrains the peptide backbone, forcing a kink in the loop conformation. Removing Pro8 relieved this constraint, allowing the CDRH3 loop to adopt a more flexible, extended conformation visible in the ESMFold prediction. This conformational shift improved the composite computational score from 0.9738 to 0.9810, advancing the sequence from rank 17 to rank 3 out of 29,574 total candidates.</p><p>In another case, parent sequence SSYTITNTQK ranked 367, while novel sequence SSTITNTQK ranked 12. A single Tyr3 deletion distinguished the parent from the novel generated sequence. Tyrosine carries a bulky aromatic ring that projects outward from the peptide backbone, imposing steric constraints on the surrounding loop geometry, which was clearly visible in the ESMFold structure of the parent. Its removal produced a flatter, more compact loop architecture, reducing steric bulk in a region predicted to be relevant for PrPSc engagement. This structural change drove the largest score improvement across all three mutation pairs. The combined score rose from 0.9439 to 0.9759, a +0.0320 improvement, and rank jumped from 367 to 12.</p><p>This study shows that machine learning can prioritize and refine antibody candidate sequences for Creutzfeldt-Jakob Disease despite extremely limited training data. Both predictive classifiers distinguished anti-prion antibodies from unrelated background sequences using only sequence-derived features. The neutralization model achieved strong discriminative performance, while the specificity classifier demonstrated moderate yet statistically significant predictive capability. These results indicate that meaningful sequence-level signal is encoded within antibody sequence features such as residue composition, charge distribution, and hydrophobicity.</p><p>Composite ranking across three independent computational screening constraints, including blood-brain barrier penetration, prion-specific binding, and neutralization potential, allowed the pipeline to approximate antibody candidate-prioritization trade-offs (Dobson et al. 2016; Mieczkowski et al. 2023). Rather than optimizing a single metric, the framework balances multiple biological requirements simultaneously, reflecting the complexity of CNS antibody candidate development. This multi-objective scoring approach enabled identification of candidates that perform well across all criteria rather than excelling in only one dimension.</p><p>External control comparison using FDA-approved CNS antibodies provided additional evidence that the models capture prion-specific sequence patterns rather than generic properties of antibodies capable of entering the brain. These control antibodies scored near zero in both predictive models, while anti-prion exemplars clustered near the highest score ranges. This separation supports that the classifiers learned biologically relevant signal associated with prion targeting rather than simply recognizing general CNS antibody traits.</p><p>The mutation model further expanded the search space by generating targeted sequence modifications around top-ranking candidates. Guided amino acid substitutions preserved biochemical compatibility while introducing structural diversity, enabling exploration of nearby sequence space. Several mutation-derived variants achieved scores comparable to or exceeding their parent sequences, demonstrating the potential for computational prioritization of candidate antibody sequences.</p><p>Traditional anti-prion drug discovery is bottlenecked by the scarcity of validated antibody datasets and the difficulty of experimental prion screening. This pipeline demonstrates that sequence-level machine learning can extract meaningful biological signal from as few as 21 positive examples, but the findings remain preliminary and computational. The multi-constraint scoring framework is disease-agnostic, and the same architecture could be retrained for Alzheimer’s, Parkinson’s, and other protein misfolding disorders where BBB penetration and isoform selectivity are simultaneously required. The identified candidates represent a prioritized set for downstream validation, including molecular docking against PrPSc crystal structures, binding affinity assays, cell-based BBB transcytosis models, and prion cell model neutralization testing. Future work would center around gathering more positive exemplars, wet-lab validation, expanding the mutation model, and replacing the BBB proxy with a trained model.</p>","references":[{"reference":"Aguzzi A, Calella AM. 2009. Prions: Protein aggregation and infectious diseases. Physiological Reviews. 89: 1105.","pubmedId":"","doi":"10.1152/physrev.00006.2009"},{"reference":"Dobson CL, Devine PW, Phillips JJ, Higazi DR, Lloyd C, Popplewell AG. 2016. Engineering the surface properties of a human monoclonal antibody prevents self-association and rapid clearance in vivo. Scientific Reports. 6: 38644.","pubmedId":"","doi":"10.1038/srep38644"},{"reference":"Mieczkowski C, Zhang X, Lee D, Nguyen K, Lv W, Wang Y, et al., Gries JM. 2023. Blueprint for antibody biologics developability. mAbs. 15: 2185924.","pubmedId":"","doi":"10.1080/19420862.2023.2185924"},{"reference":"Pankiewicz JE, Sanchez S, Kirshenbaum K, Kascsak RB, Kascsak RJ, Sadowski MJ. 2019. Anti-prion protein antibody 6D11 restores cellular proteostasis of prion protein through disrupting recycling propagation of PrPSc and targeting PrPSc for lysosomal degradation. Molecular Neurobiology. 56: 2073.","pubmedId":"","doi":"10.1007/s12035-018-1208-4"},{"reference":"Prusiner SB. 1998. Prions. Proceedings of the National Academy of Sciences. 95: 13363.","pubmedId":"","doi":"10.1073/pnas.95.23.13363"},{"reference":"Ruiz Lopez E. 2021. Transportation of single-domain antibodies through the blood-brain barrier. Biomolecules. 11: 1131.","pubmedId":"","doi":"10.3390/biom11081131"},{"reference":"Triguero D, Buciak JB, Pardridge WM. 1989. Blood-brain barrier transport of cationized immunoglobulin G. Proceedings of the National Academy of Sciences. 86: 4761.","pubmedId":"","doi":"10.1073/pnas.86.12.4761"},{"reference":"Uger MD, Chai V, Ciolfi V, Cashman NR, Tian B, Wong WY, Chao HLM. 2016. Antibodies and conjugates that target misfolded prion protein.","pubmedId":"","doi":""},{"reference":"Williamson RA, Burton DR, Prusiner SB. 2001. Antibodies specific for native PrPSc.","pubmedId":"","doi":""},{"reference":"Yang W. 2014. Efficient penetration of the blood-brain barrier by anti-prion antibody fragments. Science Translational Medicine. 6: 234ra60.","pubmedId":"","doi":"10.1126/scitranslmed.3008870"},{"reference":"Collinge J, Hawke S. 2009. Prion inhibition.","pubmedId":"","doi":""}],"title":"<p>Against All Odds: Computational Screening Via Machine Learning Ranking and Generation of Antibody Candidates for Creutzfeldt-Jakob Disease</p>","reviews":[{"reviewer":{"displayName":"Robert Mercer"},"openAcknowledgement":false,"status":{"submitted":true}}],"curatorReviews":[]},{"id":"b36e3284-009d-475c-bb30-f70e57e5ae3f","decision":"accept","abstract":"<p>Creutzfeldt-Jakob disease (CJD) is a fatal prion disorder with no approved treatments. This study developed a machine learning pipeline trained on 21 literature-curated anti-prion CDR sequences to rank 29,574 original and mutation-generated candidate sequences using neutralization, selectivity, and literature-based blood-brain barrier (BBB) proxy scores. The models showed internal performance (neutralization AUC = 0.9239; selectivity AUC = 0.7759), identified 10 high-percentile candidates, and generated 7,309 novel variants. These findings support hypothesis-generating sequence-level prioritization of anti-prion candidates, but do not establish native PrPSc-specific binding, full-antibody efficacy, exact PrP epitope recognition, or in vivo BBB penetration and require further experimental validation and testing.</p>","acknowledgements":"<p>The authors would like to thank Academy of the Canyons for providing resources to enable this project.</p>","authors":[{"affiliations":["College of the Canyons, Santa Clarita, California, United States","Academy of the Canyons, Santa Clarita, California, United States"],"departments":["",""],"credit":["conceptualization","dataCuration","formalAnalysis","fundingAcquisition","investigation","methodology","project","resources","software","validation","visualization","writing_originalDraft","writing_reviewEditing"],"email":"vtiruvellore@my.canyons.edu","firstName":"Vinayan","lastName":"Tiruvellore","submittingAuthor":true,"correspondingAuthor":true,"equalContribution":false,"WBId":null,"orcid":"https://orcid.org/0009-0003-5451-2503"},{"affiliations":["Academy of the Canyons, Santa Clarita, California, United States","College of the Canyons, Santa Clarita, California, United States"],"departments":["",""],"credit":["supervision","resources"],"email":"mkoegle@hartdistrict.org","firstName":"Mike","lastName":"Koegle","submittingAuthor":false,"correspondingAuthor":false,"equalContribution":false,"WBId":null,"orcid":null}],"awards":[],"conflictsOfInterest":"<p>The authors declare that there are no conflicts of interest present.</p>","dataTable":{"url":null},"extendedData":[],"funding":"<p>This work did not receive external funding. Support was provided by the authors and Academy of the Canyons, Santa Clarita, CA.</p>","image":{"url":"https://portal.micropublication.org/uploads/704d795891ef24bc74c9e5699dd42d7f.png"},"imageCaption":"<p>A) Machine learning pipeline for sequence-level computational prioritization of anti-prion CDR candidates. flow chart overview of the computational workflow used to evaluate CDR sequences against three screening criteria: blood-brain barrier compatibility proxy, selectivity-like scoring, and predicted neutralization-like activity. Background antibody sequences were scored relative to  anti-prion exemplars, ranked by composite candidate-prioritization score, and used as the basis for mutation-based generation of novel candidate variants.</p><p>B) Feature engineering and sequence-level model inputs. Quantitative physicochemical descriptors were extracted from antibody CDR sequences, including length, hydropathy, residue composition, charge distribution, aromatic fraction, and glycine/proline content. These engineered sequence features were used to train classifiers that distinguish anti-prion exemplar sequences from non-prion background sequences and support multi-objective ranking.</p><p>C) Model performance and candidate ranking results. Predicted neutralization-like and selectivity-like scores separated curated anti-prion positive exemplars from background antibodies and FDA-approved CNS antibody controls. ROC AUC and Mann-Whitney U test results support statistically significant discrimination between anti-prion exemplar sequences and non-prion comparison groups.</p><p>D) Structural comparison of parent and mutation-derived antibody candidates. ESMFold-predicted CDR structures show conformational effects of single-residue deletions in high-scoring mutation-derived variants. Removal of structurally restrictive residues altered loop geometry and corresponded with improved composite computational ranking relative to the parent sequences.</p>","imageTitle":"<p><i>Fig. 1 (A, B, C, D)</i></p>","methods":"<p>Data collection: This study was designed as a preliminary computational screening pipeline for candidate anti-prion CDR sequences rather than as evidence of therapeutic efficacy. The positive set consisted of 21 curated anti-prion CDR sequences compiled from patents and literature-derived criteria sets: 8 sequences in the neutralization set, 11 in the selectivity set, and 2 BBB-related comparator sequences. The modeled sequences represent variable-length CDR segments ranging from 5 to 17 amino acids and therefore capture only CDR-derived subregions of the antibody variable domain rather than a complete antibody molecule. Because experimentally validated anti-prion antibody datasets are scarce, this curated set was used as a small hypothesis-generating positive reference set. For model contrast, more than 20,000 non-prion human IGH background CDR sequences were used as negatives during training. Additionally, the source reference, sequence identity, exact sequence used, region, and model role for each exemplar are provided in exemplar_table_full.csv in the project GitHub repository.</p><p>Control construction: The primary negative control was the non-prion IGH background set. An external comparator set consisting of 8 FDA-approved CNS-related antibodies that were not designed to bind prions was used during downstream evaluation. These FDA CNS antibodies were treated as non-prion comparators to test whether the scoring framework was capturing prion-related sequence patterns rather than general CNS-targeting characteristics.</p><p>Feature engineering: Each CDR sequence was converted into a hybrid feature representation consisting of character-level 3-mer sequence features and 7 physicochemical features. The physicochemical features were length, hydrophobic fraction, aromatic fraction, positive-residue fraction, negative-residue fraction, net charge per residue, and glycine/proline fraction. In the final target feature matrix, the 21 positive sequences were represented by 137 distinct 3-mer features and 7 physicochemical features. This representation was chosen to preserve short sequence motif information while also capturing coarse biochemical properties relevant to antibody recognition and developability.</p><p>Machine-learning approach: Two binary classifiers were trained: a neutralization-like classifier and a selectivity-like classifier. Both models used L2-regularized logistic regression with class balancing and the liblinear solver. Positive labels were assigned from the curated anti-prion sequence set, and negative labels were assigned from the non-prion background set. The models therefore estimated whether a sequence more closely resembled the curated neutralization or selectivity exemplars than the background repertoire. The computational framework was trained on sequence-derived features from curated anti-prion CDR exemplars and was not trained directly against the full human PrPC sequence. It does not predict a specific PrP binding epitope. These outputs should therefore be interpreted as ranking scores rather than direct predictions of full-antibody therapeutic activity or native PrPSc-specific binding.</p><p>Training and validation: Model performance was evaluated using 5-fold stratified cross-validation with ROC AUC and average precision. The neutralization-like classifier achieved mean ROC AUC = 0.9239 ± 0.0810, and the selectivity-like classifier achieved mean ROC AUC = 0.7759 ± 0.1217 across folds. Because only 21 curated positive exemplars were available, these models should be interpreted as preliminary ranking models rather than definitive predictive classifiers. The larger variance observed for the selectivity model is consistent with reduced stability under limited positive-sample conditions.</p><p>BBB proxy scoring model: The BBB proxy scoring model was a hand-engineered sequence-ranking function rather than an experimentally trained BBB penetration model. For each antibody sequence, the model computed six physicochemical descriptors: sequence length, approximate molecular weight, net charge at pH 7.4, estimated isoelectric point (pI), GRAVY hydropathy (Kyte-Doolittle average hydropathy), and a hydrophobic patch score defined as the fraction of 6-residue windows containing at least 4 hydrophobic residues. The proxy then assigned a BBB compatibility score by rewarding sequences with net charge near 6.0 and pI near 9.6 using Gaussian-like soft-band terms, while penalizing excessive charge, excessively high pI, large molecular weight, high hydropathy, and high hydrophobic patchiness using smooth soft-penalty terms. The raw score was computed as:</p><p>0.35 x charge_reward + 0.25 x pI_reward - 0.15 x size_penalty - 0.15 x gravy_penalty - 0.10 x patch_penalty - 0.10 x charge_extreme_penalty - 0.10 x pI_extreme_penalty</p><p>and then transformed to a 0-1 scale with the logistic function where BBB_score = 1 / (1 + e<sup>(-4 * (raw - 0.3)</sup>)). Higher scores indicated stronger predicted BBB compatibility under this proxy, whereas lower scores indicated weaker compatibility. Because no independent BBB transcytosis or transport training dataset was used, this score should be interpreted strictly as a computational ranking proxy for sequence-level BBB compatibility rather than as direct evidence that intact antibodies would penetrate the BBB in vivo.</p><p>Statistical adequacy and overfitting: The positive reference set was small, so overfitting risk is an important limitation. To mitigate this, the classifiers used regularized logistic regression, class balancing, and 5-fold stratified cross-validation, and were evaluated against both a large background repertoire and an external FDA CNS comparator set. These steps support the use of the framework for preliminary prioritization, but they do not eliminate the need for independent experimental validation. Accordingly, all outputs should be interpreted as hypothesis-generating candidate rankings rather than validated therapeutic predictions.</p><p>Candidate ranking: Candidate sequences were ranked by combining the neutralization-like score, the selectivity-like score, and the BBB proxy score. Ranking was performed on original seed sequences and mutation-generated variants. Across 29,574 evaluated sequences in the ranking/validation workflow, the top candidates were prioritized based on combined percentile performance across the three criteria.</p><p>Statistical testing: Mann-Whitney U tests were used to compare score distributions between predefined groups, including positives versus background and positives versus FDA CNS comparators, for both neutralization-like and selectivity-like scores. These tests were used as score-separation analyses rather than proof of causal biological discrimination or native PrPSc-specific binding.</p><p>Mutation modeling: Sequence diversification was performed computationally by targeted mutation and rescoring. This procedure generated 7,309 novel sequences not present in the original library. These mutated candidates were evaluated using the same ranking framework as the original sequences.</p><p>Data and code availability: The curated input datasets, extracted feature tables, trained scoring artifacts, ranked candidate outputs, analysis code, and the exemplar table with the CDR sequences the model was trained on are publicly available at https://github.com/Axy-lis/CJD_ML_AAO.</p>","reagents":"<p></p>","patternDescription":"<p>CJD is a rare but universally fatal prion disease. There is no cure, no disease-modifying therapy. CJD is caused by the misfolding of native prion protein, PrPC, into the pathogenic isoform PrPSc. PrPSc acts as a self-propagating template, converts normal proteins into misfolded forms, forms amyloid aggregates resistant to degradation, and causes widespread neuronal death (Prusiner 1998; Aguzzi and Calella 2009). Therapeutic development is difficult because PrPSc is structurally similar to PrPC, the central nervous system is protected by the highly selective BBB, candidate interventions must avoid disrupting normal prion protein function, and despite decades of research, no clinically approved antibody or small-molecule therapy exists (Aguzzi and Calella 2009; Pankiewicz et al. 2019).</p><p>Many reported anti-prion antibodies act through binding to PrPC or related conformational states, and apparent PrPSc recognition can depend on denaturation or assay-specific exposure of epitopes. Accordingly, the present study should not be interpreted as establishing selective binding to native PrPSc. In addition, prior studies have reported neurotoxicity associated with some anti-prion antibodies and related PrP toxic signaling pathways, emphasizing that antibody-based intervention in prion disease remains biologically complex and that the present computational rankings should not be interpreted as evidence of safety or therapeutic suitability (Sonati et al. 2013; Reimann et al. 2016; Wu et al. 2017; Frontzek et al. 2022; Mercer and Harris 2023).</p><p>Machine learning approaches to antibody engineering are promising in identifying candidate antibodies. By training ranking models on sequence-derived CDR features from curated anti-prion antibody exemplars, it is possible to evaluate large candidate libraries and identify those most likely to satisfy computational screening criteria. The experimental question was whether a machine learning-driven computational pipeline could prioritize candidate anti-prion CDR segments for CJD using neutralization-like, selectivity-like, and BBB compatibility proxy criteria. The hypothesis was that a machine learning pipeline trained on sequence-derived features from literature-curated anti-prion CDR exemplars would rank anti-prion-like candidates above background sequences, distinguish them from unrelated CNS antibodies, and generate novel, high-scoring variants through mutation modeling. The modeled inputs in this study were variable-length CDR segments rather than complete antibody molecules or fixed-length 10-mers. In the curated exemplar set, these CDR sequences ranged from 5 to 17 amino acids, with many clustering near approximately 10 amino acids. The framework was therefore designed for sequence-level prioritization rather than full-antibody functional prediction.</p><p>The primary control was baseline background antibodies. An external negative control used FDA-approved CNS-targeting antibodies to test whether the model captured anti-prion sequence signatures rather than general CNS-targeting characteristics. For data collection, 21 literature-curated anti-prion CDR exemplars were used as positives, and more than 20,000 non-prion sequences were used as background. Anti-prion antibody examples and prion-targeting antibody concepts were drawn from prior literature and patent sources (Pankiewicz et al. 2019; Uger et al. 2016; Williamson et al. 2001; Collinge and Hawke 2009). Sequences derived from Uger et al. 2016 were included as literature-reported anti-prion exemplars for computational comparison, but not as definitive evidence of selective native PrPSc recognition. For feature engineering, CDR sequences were converted into quantitative descriptors, including 3-mer motif embeddings, amino acid composition, length, net charge at pH 7.4, hydropathy using the Kyte-Doolittle scale, aromatic fraction, and residue class proportions.</p><p>Model training was designed to distinguish positives from background. The pipeline included a neutralization-like classifier, a selectivity-like classifier, and a BBB compatibility proxy. Candidate ranking was performed by scoring more than 20,000 sequences and calculating a composite score across all three criteria. The top 25 candidates were selected. Mutation modeling was then applied through targeted sequence modifications, and modified sequences were re-scored to identify novel high-ranking candidates.</p><p>The neutralization-like classifier achieved ROC AUC = 0.9239 ± 0.0810, and the selectivity-like classifier achieved ROC AUC = 0.7759 ± 0.1217. Both were above random expectation, supporting that CDR sequence-derived features carry meaningful signal for anti-prion exemplar ranking. The BBB compatibility proxy was a literature-based heuristic informed by papers on antibody BBB transport and anti-prion antibody fragments (Triguero et al. 1989; Yang et al. 2014; Ruiz-López et al. 2021). Strong internal model performance was achieved despite training on only 21 curated exemplars.</p><p>Mann-Whitney U tests showed statistically significant score separation between positive exemplars and background sequences and between positives and FDA CNS antibodies. Positives versus background for neutralization had a U statistic of 180000 and a p-value of 0.000000102. Positives versus background for selectivity had a U statistic of 300000 and a p-value of 0.0000000001. Positives versus FDA CNS antibodies for neutralization had a U statistic of 72 and a p-value of 0.0003129. Positives versus FDA CNS antibodies for selectivity had a U statistic of 120 and a p-value of 0.0000609. These low p-values support score separation between predefined groups, but do not independently establish antibody activity, selective native PrPSc recognition, or guarantee model generalizability.</p><p>The neutralization score distributions showed positive exemplars clustered near the highest predicted values, around 0.99 to 1.0, whereas background sequences remained spread across lower score values, around 0.3 to 0.7. The selectivity-like score distribution showed a similar pattern, with positive exemplars again occupying the highest scoring region and background sequences occupying lower score ranges. Neutralization-like and selectivity-like scores were also compared across known anti-prion exemplars, FDA CNS antibodies, and background sequences. Positive exemplars clustered near maximal scores, FDA CNS antibodies scored near zero, and background sequences showed low-score distributions. Together, these comparisons support that the classifiers captured anti-prion sequence signatures rather than general antibody features.</p><p>The ranked candidate pool consisted of more than 20,000 original antibody sequences and about 7,000 mutation-generated variants. This search space allowed the machine learning models to identify rare high-scoring candidates within diverse antibody sequence space. Ten candidates ranked in the 95th to 99th percentile across 29,574 evaluated sequences. Positive exemplars were heavily enriched in the top 50, while FDA CNS antibodies were entirely absent. Mutation modeling generated 7,309 novel sequences not present in the original library. The top novel candidates ranked 3rd, 7th, and 12th overall, outperforming 99.97% of all evaluated sequences in this computational screen.</p><p>Single-residue deletions produced the largest improvements. In one case, parent sequence QQNKNWPPGT ranked 17, while novel sequence QQNKNWPGT ranked 3. A single Pro8 deletion distinguished the parent from the novel generated sequence. Removing Pro8 relieved local conformational constraint in the CDR-segment loop and improved the composite computational score from 0.9738 to 0.9810, advancing the sequence from rank 17 to rank 3 out of 29,574 total candidates. In another case, parent sequence SSYTITNTQK ranked 367, while novel sequence SSTITNTQK ranked 12. A single Tyr3 deletion distinguished the parent from the novel generated sequence and increased the combined score from 0.9439 to 0.9759, producing the largest score improvement across the highlighted mutation pairs. These examples suggest that small local sequence changes can measurably alter the ranking outputs of the computational framework.</p><p>This study shows that machine learning can prioritize and refine anti-prion CDR candidate sequences for Creutzfeldt-Jakob disease despite extremely limited training data. Both predictive classifiers distinguished anti-prion exemplars from unrelated background sequences using only sequence-derived features. The neutralization-like model achieved strong discriminative performance, while the selectivity-like classifier demonstrated moderate yet statistically significant predictive capability. These results indicate that meaningful sequence-level signal is encoded within antibody sequence features such as residue composition, charge distribution, hydrophobicity, and short motif content.</p><p>Composite ranking across three independent computational screening constraints, including BBB compatibility proxy, anti-prion selectivity-like scoring, and neutralization potential, allowed the pipeline to approximate candidate-prioritization trade-offs (Dobson et al. 2016; Mieczkowski et al. 2023). Rather than optimizing a single metric, the framework balances multiple biological requirements simultaneously, reflecting the complexity of CNS antibody candidate development. This multi-objective scoring approach enabled identification of candidates that perform well across all criteria rather than excelling in only one dimension.</p><p>External control comparison using FDA-approved CNS antibodies provided additional evidence that the models capture anti-prion sequence patterns rather than generic CNS-associated properties. These control antibodies scored near zero in both predictive models, while anti-prion exemplars clustered near the highest score ranges. This separation supports that the classifiers learned biologically relevant signal associated with prion targeting rather than simply recognizing general CNS antibody traits.</p><p>The mutation model further expanded the search space by generating targeted sequence modifications around top-ranking candidates. Several mutation-derived variants achieved scores comparable to or exceeding their parent sequences, demonstrating the potential for computational prioritization of candidate antibody sequences. Traditional anti-prion drug discovery is bottlenecked by the scarcity of validated antibody datasets and the difficulty of experimental prion screening. This pipeline demonstrates that sequence-level machine learning can extract meaningful biological signal from as few as 21 positive examples, but the findings remain preliminary and computational. The identified candidates represent a prioritized set for later validation, including molecular docking against PrP structures, binding affinity assays, cell-based BBB transcytosis models, and prion cell model neutralization testing. Future work would center around gathering more positive exemplars, wet-lab validation, expanding the mutation model, and replacing the BBB proxy with a trained transport model. These results should be interpreted as hypothesis for sequence ranking rather than proof of native PrPSc-specific binding, defined PrP epitope recognition, or BBB permeation. These results should also be interpreted in light of prior reports of anti-prion antibody-associated neurotoxicity.</p>","references":[{"reference":"Aguzzi A, Calella AM. 2009. Prions: Protein aggregation and infectious diseases. Physiological Reviews. 89: 1105.","pubmedId":"","doi":"10.1152/physrev.00006.2009"},{"reference":"Dobson CL, Devine PW, Phillips JJ, Higazi DR, Lloyd C, Popplewell AG. 2016. Engineering the surface properties of a human monoclonal antibody prevents self-association and rapid clearance in vivo. Scientific Reports. 6: 38644.","pubmedId":"","doi":"10.1038/srep38644"},{"reference":"Mieczkowski C, Zhang X, Lee D, Nguyen K, Lv W, Wang Y, et al., Gries JM. 2023. Blueprint for antibody biologics developability. mAbs. 15: 2185924.","pubmedId":"","doi":"10.1080/19420862.2023.2185924"},{"reference":"Pankiewicz JE, Sanchez S, Kirshenbaum K, Kascsak RB, Kascsak RJ, Sadowski MJ. 2019. Anti-prion protein antibody 6D11 restores cellular proteostasis of prion protein through disrupting recycling propagation of PrPSc and targeting PrPSc for lysosomal degradation. Molecular Neurobiology. 56: 2073.","pubmedId":"","doi":"10.1007/s12035-018-1208-4"},{"reference":"Prusiner SB. 1998. Prions. Proceedings of the National Academy of Sciences. 95: 13363.","pubmedId":"","doi":"10.1073/pnas.95.23.13363"},{"reference":"<p>Ruiz-López E, Schuhmacher AJ. 2021. Transportation of Single-Domain Antibodies through the Blood-Brain Barrier. Biomolecules 11(8): 10.3390/biom11081131.</p>","pubmedId":"34439797","doi":""},{"reference":"Triguero D, Buciak JB, Pardridge WM. 1989. Blood-brain barrier transport of cationized immunoglobulin G. Proceedings of the National Academy of Sciences. 86: 4761.","pubmedId":"","doi":"10.1073/pnas.86.12.4761"},{"reference":"Uger MD, Chai V, Ciolfi V, Cashman NR, Tian B, Wong WY, Chao HLM. 2016. Antibodies and conjugates that target misfolded prion protein.","pubmedId":"","doi":""},{"reference":"Williamson RA, Burton DR, Prusiner SB. 2001. Antibodies specific for native PrPSc.","pubmedId":"","doi":""},{"reference":"Yang W. 2014. Efficient penetration of the blood-brain barrier by anti-prion antibody fragments. Science Translational Medicine. 6: 234ra60.","pubmedId":"","doi":"10.1126/scitranslmed.3008870"},{"reference":"Collinge J, Hawke S. 2009. Prion inhibition.","pubmedId":"","doi":""},{"reference":"<p>Sonati T, Reimann RR, Falsig J, Baral PK, O'Connor T, Hornemann S, et al., Aguzzi A. 2013. The toxicity of antiprion antibodies is mediated by the flexible tail of the prion protein. Nature 501(7465): 102-6.</p>","pubmedId":"23903654","doi":""},{"reference":"<p>Reimann RR, Sonati T, Hornemann S, Herrmann US, Arand M, Hawke S, Aguzzi A. 2016. Differential Toxicity of Antibodies to the Prion Protein. PLoS Pathog 12(1): e1005401.</p>","pubmedId":"26821311","doi":""},{"reference":"<p>Wu B, McDonald AJ, Markham K, Rich CB, McHugh KP, Tatzelt J, et al., Harris DA. 2017. The N-terminus of the prion protein is a toxic effector regulated by the C-terminus. Elife 6: pii: e23473. 10.7554/eLife.23473.</p>","pubmedId":"28527237","doi":""},{"reference":"<p>Frontzek K, Bardelli M, Senatore A, Henzi A, Reimann RR, Bedir S, et al., Aguzzi A. 2022. A conformational switch controlling the toxicity of the prion protein. Nat Struct Mol Biol 29(8): 831-840.</p>","pubmedId":"35948768","doi":""},{"reference":"<p>Mercer RCC, Harris DA. 2023. Mechanisms of prion-induced toxicity. Cell Tissue Res 392(1): 81-96.</p>","pubmedId":"36070155","doi":""}],"title":"<p>Against All Odds: Computational Screening Via Machine Learning Ranking and Generation of Antibody Candidates for Creutzfeldt-Jakob Disease</p>","reviews":[{"reviewer":{"displayName":"Robert Mercer"},"openAcknowledgement":false,"status":{"submitted":true}}],"curatorReviews":[]},{"id":"05f963e4-2cf9-4dfe-ade5-df32accc957f","decision":"edit","abstract":"<p>Creutzfeldt-Jakob disease (CJD) is a fatal prion disorder with no approved treatments. This study developed a machine learning pipeline trained on 21 literature-curated PrP-targeting CDR sequences to rank 29,574 original and mutation-generated candidate sequences using neutralization, selectivity, and literature-based blood-brain barrier (BBB) proxy scores. The models showed internal performance (neutralization AUC = 0.9239; selectivity AUC = 0.7759), identified 10 high-percentile candidates, and generated 7,309 novel variants. These findings support hypothesis-generating sequence-level prioritization of PrP-targeting candidates, but do not establish native PrP<sup>Sc</sup>-specific binding, full-antibody efficacy, exact PrP epitope recognition, or in vivo BBB penetration and require further experimental validation and testing.</p>","acknowledgements":"<p>The authors would like to thank Academy of the Canyons for providing resources to enable this project.</p>","authors":[{"affiliations":["College of the Canyons, Santa Clarita, California, United States","Academy of the Canyons, Santa Clarita, California, United States"],"departments":["",""],"credit":["conceptualization","dataCuration","formalAnalysis","fundingAcquisition","investigation","methodology","project","resources","software","validation","visualization","writing_originalDraft","writing_reviewEditing"],"email":"vtiruvellore@my.canyons.edu","firstName":"Vinayan","lastName":"Tiruvellore","submittingAuthor":true,"correspondingAuthor":true,"equalContribution":false,"WBId":null,"orcid":"https://orcid.org/0009-0003-5451-2503"},{"affiliations":["Academy of the Canyons, Santa Clarita, California, United States","College of the Canyons, Santa Clarita, California, United States"],"departments":["",""],"credit":["supervision","resources"],"email":"mkoegle@hartdistrict.org","firstName":"Mike","lastName":"Koegle","submittingAuthor":false,"correspondingAuthor":false,"equalContribution":false,"WBId":null,"orcid":null}],"awards":[],"conflictsOfInterest":"<p>The authors declare that there are no conflicts of interest present.</p>","dataTable":{"url":null},"extendedData":[],"funding":"<p>This work did not receive external funding. Support was provided by the authors and Academy of the Canyons, Santa Clarita, CA.</p>","image":{"url":"https://portal.micropublication.org/uploads/704d795891ef24bc74c9e5699dd42d7f.png"},"imageCaption":"<p>A) Machine learning pipeline for sequence-level computational prioritization of PrP-targeting CDR candidates. flow chart overview of the computational workflow used to evaluate CDR sequences against three screening criteria: blood-brain barrier compatibility proxy, selectivity-like scoring, and predicted neutralization-like activity. Background antibody sequences were scored relative to PrP-targeting exemplars, ranked by composite candidate-prioritization score, and used as the basis for mutation-based generation of novel candidate variants.</p><p>B) Feature engineering and sequence-level model inputs. Quantitative physicochemical descriptors were extracted from antibody CDR sequences, including length, hydropathy, residue composition, charge distribution, aromatic fraction, and glycine/proline content. These engineered sequence features were used to train classifiers that distinguish PrP-targeting exemplar sequences from non-PrP background sequences and support multi-objective ranking.</p><p>C) Model performance and candidate ranking results. Predicted neutralization-like and selectivity-like scores separated curated PrP-targeting positive exemplars from background antibodies and FDA-approved CNS antibody controls. ROC AUC and Mann-Whitney U test results support statistically significant discrimination between PrP-targeting exemplar sequences and non-PrP comparison groups.</p><p>D) Structural comparison of parent and mutation-derived antibody candidates. ESMFold-predicted CDR structures show conformational effects of single-residue deletions in high-scoring mutation-derived variants. Removal of structurally restrictive residues altered loop geometry and corresponded with improved composite computational ranking relative to the parent sequences.</p>","imageTitle":"<p><i>Fig. 1 (A, B, C, D)</i></p>","methods":"<p>Data collection: This study was designed as a preliminary computational screening pipeline for candidate PrP-targeting CDR sequences rather than as evidence of therapeutic efficacy. The positive set consisted of 21 curated PrP-targeting CDR sequences compiled from patents and literature-derived criteria sets: 8 sequences in the neutralization set, 11 in the selectivity set, and 2 BBB-related comparator sequences. The modeled sequences represent variable-length CDR segments ranging from 5 to 17 amino acids and therefore capture only CDR-derived subregions of the antibody variable domain rather than a complete antibody molecule. Because experimentally validated anti-PrP antibody datasets are scarce, this curated set was used as a small hypothesis-generating positive reference set. For model contrast, more than 20,000 non-PrP human IGH background CDR sequences were used as negatives during training. Additionally, the source reference, sequence identity, exact sequence used, region, and model role for each exemplar are provided in exemplar_table_full.csv in the project GitHub repository.</p><p>Control construction: The primary negative control was the non-PrP IGH background set. An external comparator set consisting of 8 FDA-approved CNS-related antibodies that were not designed to bind PrP was used during downstream evaluation. These FDA CNS antibodies were treated as non-PrP comparators to test whether the scoring framework was capturing PrP-related sequence patterns rather than general CNS-targeting characteristics.</p><p>Feature engineering: Each CDR sequence was converted into a hybrid feature representation consisting of character-level 3-mer sequence features and 7 physicochemical features. The physicochemical features were length, hydrophobic fraction, aromatic fraction, positive-residue fraction, negative-residue fraction, net charge per residue, and glycine/proline fraction. In the final target feature matrix, the 21 positive sequences were represented by 137 distinct 3-mer features and 7 physicochemical features. This representation was chosen to preserve short sequence motif information while also capturing coarse biochemical properties relevant to antibody recognition and developability.</p><p>Machine-learning approach: Two binary classifiers were trained: a neutralization-like classifier and a selectivity-like classifier. Both models used L2-regularized logistic regression with class balancing and the liblinear solver. Positive labels were assigned from the curated PrP-targeting sequence set, and negative labels were assigned from the non-PrP background set. The models therefore estimated whether a sequence more closely resembled the curated neutralization or selectivity exemplars than the background repertoire. The computational framework was trained on sequence-derived features from curated PrP-targeting CDR exemplars and was not trained directly against the full human PrP<sup>C</sup> sequence. It does not predict a specific PrP binding epitope. These outputs should therefore be interpreted as ranking scores rather than direct predictions of full-antibody therapeutic activity or native PrP<sup>Sc</sup>-specific binding.</p><p>Training and validation: Model performance was evaluated using 5-fold stratified cross-validation with ROC AUC and average precision. The neutralization-like classifier achieved mean ROC AUC = 0.9239 ± 0.0810, and the selectivity-like classifier achieved mean ROC AUC = 0.7759 ± 0.1217 across folds. Because only 21 curated positive exemplars were available, these models should be interpreted as preliminary ranking models rather than definitive predictive classifiers. The larger variance observed for the selectivity model is consistent with reduced stability under limited positive-sample conditions.</p><p>BBB proxy scoring model: The BBB proxy scoring model was a hand-engineered sequence-ranking function rather than an experimentally trained BBB penetration model. For each antibody sequence, the model computed six physicochemical descriptors: sequence length, approximate molecular weight, net charge at pH 7.4, estimated isoelectric point (pI), GRAVY hydropathy (Kyte-Doolittle average hydropathy), and a hydrophobic patch score defined as the fraction of 6-residue windows containing at least 4 hydrophobic residues. The proxy then assigned a BBB compatibility score by rewarding sequences with net charge near 6.0 and pI near 9.6 using Gaussian-like soft-band terms, while penalizing excessive charge, excessively high pI, large molecular weight, high hydropathy, and high hydrophobic patchiness using smooth soft-penalty terms. The raw score was computed as:</p><p>0.35 x charge_reward + 0.25 x pI_reward - 0.15 x size_penalty - 0.15 x gravy_penalty - 0.10 x patch_penalty - 0.10 x charge_extreme_penalty - 0.10 x pI_extreme_penalty</p><p>and then transformed to a 0-1 scale with the logistic function where BBB_score = 1 / (1 + e<sup>(-4 * (raw - 0.3)</sup>)). Higher scores indicated stronger predicted BBB compatibility under this proxy, whereas lower scores indicated weaker compatibility. Because no independent BBB transcytosis or transport training dataset was used, this score should be interpreted strictly as a computational ranking proxy for sequence-level BBB compatibility rather than as direct evidence that intact antibodies would penetrate the BBB in vivo.</p><p>Statistical adequacy and overfitting: The positive reference set was small, so overfitting risk is an important limitation. To mitigate this, the classifiers used regularized logistic regression, class balancing, and 5-fold stratified cross-validation, and were evaluated against both a large background repertoire and an external FDA CNS comparator set. These steps support the use of the framework for preliminary prioritization, but they do not eliminate the need for independent experimental validation. Accordingly, all outputs should be interpreted as hypothesis-generating candidate rankings rather than validated therapeutic predictions.</p><p>Candidate ranking: Candidate sequences were ranked by combining the neutralization-like score, the selectivity-like score, and the BBB proxy score. Ranking was performed on original seed sequences and mutation-generated variants. Across 29,574 evaluated sequences in the ranking/validation workflow, the top candidates were prioritized based on combined percentile performance across the three criteria.</p><p>Statistical testing: Mann-Whitney U tests were used to compare score distributions between predefined groups, including positives versus background and positives versus FDA CNS comparators, for both neutralization-like and selectivity-like scores. These tests were used as score-separation analyses rather than proof of causal biological discrimination or native PrP<sup>Sc</sup>-specific binding.</p><p>Mutation modeling: Sequence diversification was performed computationally by targeted mutation and rescoring. This procedure generated 7,309 novel sequences not present in the original library. These mutated candidates were evaluated using the same ranking framework as the original sequences.</p><p>Data and code availability: The curated input datasets, extracted feature tables, trained scoring artifacts, ranked candidate outputs, analysis code, and the exemplar table with the CDR sequences the model was trained on are publicly available at https://github.com/Axy-lis/CJD_ML_AAO.</p>","reagents":"<p></p>","patternDescription":"<p>CJD is a rare but universally fatal prion disease. There is no cure, no disease-modifying therapy. CJD is caused by the misfolding of native prion protein, PrP<sup>C</sup>, into the pathogenic isoform PrP<sup>Sc</sup>. PrP<sup>Sc</sup> acts as a self-propagating template, converts normal proteins into misfolded forms, forms amyloid aggregates, and causes widespread neuronal death (Prusiner 1998; Aguzzi and Calella 2009). Therapeutic development is challenging because the pathogenic PrP<sup>Sc</sup> fold can differ substantially between species and disease-associated prion strains, so compounds effective against mouse prions may fail against human prions. Additionally, the blood-brain barrier (BBB) limits CNS delivery, and therapies must avoid disrupting normal PrP<sup>C</sup> function. Despite decades of research, no clinically approved antibody or small-molecule therapy exists (Aguzzi and Calella 2009; Pankiewicz et al. 2019).</p><p>Many reported anti-PrP antibodies act through binding to PrP<sup>C</sup> or related conformational states, and apparent PrP<sup>Sc</sup> recognition can depend on denaturation or assay-specific exposure of epitopes. Accordingly, the present study should not be interpreted as establishing selective binding to native PrP<sup>Sc</sup>. In addition, prior studies have reported neurotoxicity associated with some anti-PrP antibodies and related PrP toxic signaling pathways, emphasizing that antibody-based intervention in prion disease remains biologically complex and that the present computational rankings should not be interpreted as evidence of safety or therapeutic suitability (Sonati et al. 2013; Reimann et al. 2016; Wu et al. 2017; Frontzek et al. 2022; Mercer and Harris 2023).</p><p>Machine learning approaches to antibody engineering are promising in identifying candidate antibodies. By training ranking models on sequence-derived CDR features from curated anti-PrP antibody exemplars, it is possible to evaluate large candidate libraries and identify those most likely to satisfy computational screening criteria. The experimental question was whether a machine learning-driven computational pipeline could prioritize candidate PrP-targeting CDR segments for CJD using neutralization-like, selectivity-like, and BBB compatibility proxy criteria. The hypothesis was that a machine learning pipeline trained on sequence-derived features from literature-curated PrP-targeting CDR exemplars would rank anti-PrP-like candidates above background sequences, distinguish them from unrelated CNS antibodies, and generate novel, high-scoring variants through mutation modeling. The modeled inputs in this study were variable-length CDR segments rather than complete antibody molecules or fixed-length 10-mers. In the curated exemplar set, these CDR sequences ranged from 5 to 17 amino acids, with many clustering near approximately 10 amino acids. The framework was therefore designed for sequence-level prioritization rather than full-antibody functional prediction.</p><p>The primary control was baseline background antibodies. An external negative control used FDA-approved CNS-targeting antibodies to test whether the model captured anti-PrP sequence signatures rather than general CNS-targeting characteristics. For data collection, 21 literature-curated PrP-targeting CDR exemplars were used as positives, and more than 20,000 non-PrP sequences were used as background. Anti-PrP antibody examples and PrP-targeting antibody concepts were drawn from prior literature and patent sources (Pankiewicz et al. 2019; Uger et al. 2016; Williamson et al. 2001; Collinge and Hawke 2009). Sequences derived from Uger et al. 2016 were included as literature-reported PrP-targeting exemplars for computational comparison, but not as definitive evidence of selective native PrP<sup>Sc</sup> recognition. For feature engineering, CDR sequences were converted into quantitative descriptors, including 3-mer motif embeddings, amino acid composition, length, net charge at pH 7.4, hydropathy using the Kyte-Doolittle scale, aromatic fraction, and residue class proportions.</p><p>Model training was designed to distinguish positives from background. The pipeline included a neutralization-like classifier, a selectivity-like classifier, and a BBB compatibility proxy. Candidate ranking was performed by scoring more than 20,000 sequences and calculating a composite score across all three criteria. The top 25 candidates were selected. Mutation modeling was then applied through targeted sequence modifications, and modified sequences were re-scored to identify novel high-ranking candidates.</p><p>The neutralization-like classifier achieved ROC AUC = 0.9239 ± 0.0810, and the selectivity-like classifier achieved ROC AUC = 0.7759 ± 0.1217. Both were above random expectation, supporting that CDR sequence-derived features carry meaningful signal for PrP-targeting exemplar ranking. The BBB compatibility proxy was a literature-based heuristic informed by papers on antibody BBB transport and anti-PrP antibody fragments (Triguero et al. 1989; Yang et al. 2014; Ruiz-López et al. 2021). Strong internal model performance was achieved despite training on only 21 curated exemplars.</p><p>Mann-Whitney U tests showed statistically significant score separation between positive exemplars and background sequences and between positives and FDA CNS antibodies. Positives versus background for neutralization had a U statistic of 180000 and a p-value of 0.000000102. Positives versus background for selectivity had a U statistic of 300000 and a p-value of 0.0000000001. Positives versus FDA CNS antibodies for neutralization had a U statistic of 72 and a p-value of 0.0003129. Positives versus FDA CNS antibodies for selectivity had a U statistic of 120 and a p-value of 0.0000609. These low p-values support score separation between predefined groups, but do not independently establish antibody activity, selective native PrP<sup>Sc</sup> recognition, or guarantee model generalizability.</p><p>The neutralization score distributions showed positive exemplars clustered near the highest predicted values, around 0.99 to 1.0, whereas background sequences remained spread across lower score values, around 0.3 to 0.7. The selectivity-like score distribution showed a similar pattern, with positive exemplars again occupying the highest scoring region and background sequences occupying lower score ranges. Neutralization-like and selectivity-like scores were also compared across known PrP-targeting exemplars, FDA CNS antibodies, and background sequences. Positive exemplars clustered near maximal scores, FDA CNS antibodies scored near zero, and background sequences showed low-score distributions. Together, these comparisons support that the classifiers captured features associated with the PrP-targeting reference set rather than general antibody features.</p><p>The ranked candidate pool consisted of more than 20,000 original antibody sequences and about 7,000 mutation-generated variants. This search space allowed the machine learning models to identify rare high-scoring candidates within diverse antibody sequence space. Ten candidates ranked in the 95th to 99th percentile across 29,574 evaluated sequences. Positive exemplars were heavily enriched in the top 50, while FDA CNS antibodies were entirely absent. Mutation modeling generated 7,309 novel sequences not present in the original library. The top novel candidates ranked 3rd, 7th, and 12th overall, outperforming 99.97% of all evaluated sequences in this computational screen.</p><p>Single-residue deletions produced the largest improvements. In one case, parent sequence QQNKNWPPGT ranked 17, while novel sequence QQNKNWPGT ranked 3. A single Pro8 deletion distinguished the parent from the novel generated sequence. Removing Pro8 relieved local conformational constraint in the CDR-segment loop and improved the composite computational score from 0.9738 to 0.9810, advancing the sequence from rank 17 to rank 3 out of 29,574 total candidates. In another case, parent sequence SSYTITNTQK ranked 367, while novel sequence SSTITNTQK ranked 12. A single Tyr3 deletion distinguished the parent from the novel generated sequence and increased the combined score from 0.9439 to 0.9759, producing the largest score improvement across the highlighted mutation pairs. These examples suggest that small local sequence changes can measurably alter the ranking outputs of the computational framework.</p><p>This study shows that machine learning can prioritize and refine candidate CDR sequences for Creutzfeldt-Jakob disease despite extremely limited training data. Both predictive classifiers distinguished PrP-targeting exemplars from unrelated background sequences using only sequence-derived features. The neutralization-like model achieved strong discriminative performance, while the selectivity-like classifier demonstrated moderate yet statistically significant predictive capability. These results indicate that meaningful sequence-level signal is encoded within antibody sequence features such as residue composition, charge distribution, hydrophobicity, and short motif content.</p><p>Composite ranking across three independent computational screening constraints, including BBB compatibility proxy, anti-PrP selectivity-like scoring, and neutralization potential, allowed the pipeline to approximate candidate-prioritization trade-offs (Dobson et al. 2016; Mieczkowski et al. 2023). Rather than optimizing a single metric, the framework balances multiple biological requirements simultaneously. This multi-objective scoring approach enabled identification of candidates that perform well across all criteria rather than excelling in only one dimension.</p><p>External control comparison using FDA-approved CNS antibodies provided additional evidence that the models capture features associated with PrP-targeting rather than generic CNS-associated properties. These control antibodies scored near zero in both predictive models, while PrP-targeting exemplars clustered near the highest score ranges. This separation supports that the classifiers learned biologically relevant signal associated with PrP targeting rather than simply recognizing general CNS antibody traits.</p><p>The mutation model further expanded the search space by generating targeted sequence modifications around top-ranking candidates. Several mutation-derived variants achieved scores comparable to or exceeding their parent sequences, demonstrating the potential for computational prioritization of candidate antibody sequences. Traditional anti-PrP drug discovery is bottlenecked by the scarcity of validated antibody datasets and the difficulty of experimental prion screening. This pipeline demonstrates that sequence-level machine learning can extract meaningful biological signal from as few as 21 positive examples, but the findings remain preliminary and computational. The identified candidates represent a prioritized set for later validation, including molecular docking against PrP structures, binding affinity assays, cell-based BBB transcytosis models, and prion cell model neutralization testing. Future work would center around gathering more positive exemplars, wet-lab validation, expanding the mutation model, and replacing the BBB proxy with a trained transport model. These results should be interpreted as hypothesis for sequence ranking rather than proof of native PrP<sup>Sc</sup>-specific binding, defined PrP epitope recognition, or BBB permeation. These results should also be interpreted in light of prior reports of anti-PrP antibody-associated neurotoxicity.</p>","references":[{"reference":"Aguzzi A, Calella AM. 2009. Prions: Protein aggregation and infectious diseases. Physiological Reviews. 89: 1105.","pubmedId":"","doi":"10.1152/physrev.00006.2009"},{"reference":"Dobson CL, Devine PW, Phillips JJ, Higazi DR, Lloyd C, Popplewell AG. 2016. Engineering the surface properties of a human monoclonal antibody prevents self-association and rapid clearance in vivo. Scientific Reports. 6: 38644.","pubmedId":"","doi":"10.1038/srep38644"},{"reference":"Mieczkowski C, Zhang X, Lee D, Nguyen K, Lv W, Wang Y, et al., Gries JM. 2023. Blueprint for antibody biologics developability. mAbs. 15: 2185924.","pubmedId":"","doi":"10.1080/19420862.2023.2185924"},{"reference":"Pankiewicz JE, Sanchez S, Kirshenbaum K, Kascsak RB, Kascsak RJ, Sadowski MJ. 2019. Anti-prion protein antibody 6D11 restores cellular proteostasis of prion protein through disrupting recycling propagation of PrPSc and targeting PrPSc for lysosomal degradation. Molecular Neurobiology. 56: 2073.","pubmedId":"","doi":"10.1007/s12035-018-1208-4"},{"reference":"Prusiner SB. 1998. Prions. Proceedings of the National Academy of Sciences. 95: 13363.","pubmedId":"","doi":"10.1073/pnas.95.23.13363"},{"reference":"<p>Ruiz-López E, Schuhmacher AJ. 2021. Transportation of Single-Domain Antibodies through the Blood-Brain Barrier. Biomolecules 11(8): 10.3390/biom11081131.</p>","pubmedId":"34439797","doi":""},{"reference":"Triguero D, Buciak JB, Pardridge WM. 1989. Blood-brain barrier transport of cationized immunoglobulin G. Proceedings of the National Academy of Sciences. 86: 4761.","pubmedId":"","doi":"10.1073/pnas.86.12.4761"},{"reference":"Uger MD, Chai V, Ciolfi V, Cashman NR, Tian B, Wong WY, Chao HLM. 2016. Antibodies and conjugates that target misfolded prion protein.","pubmedId":"","doi":""},{"reference":"Williamson RA, Burton DR, Prusiner SB. 2001. Antibodies specific for native PrPSc.","pubmedId":"","doi":""},{"reference":"Yang W. 2014. Efficient penetration of the blood-brain barrier by anti-prion antibody fragments. Science Translational Medicine. 6: 234ra60.","pubmedId":"","doi":"10.1126/scitranslmed.3008870"},{"reference":"Collinge J, Hawke S. 2009. Prion inhibition.","pubmedId":"","doi":""},{"reference":"<p>Sonati T, Reimann RR, Falsig J, Baral PK, O'Connor T, Hornemann S, et al., Aguzzi A. 2013. The toxicity of antiprion antibodies is mediated by the flexible tail of the prion protein. Nature 501(7465): 102-6.</p>","pubmedId":"23903654","doi":""},{"reference":"<p>Reimann RR, Sonati T, Hornemann S, Herrmann US, Arand M, Hawke S, Aguzzi A. 2016. Differential Toxicity of Antibodies to the Prion Protein. PLoS Pathog 12(1): e1005401.</p>","pubmedId":"26821311","doi":""},{"reference":"<p>Wu B, McDonald AJ, Markham K, Rich CB, McHugh KP, Tatzelt J, et al., Harris DA. 2017. The N-terminus of the prion protein is a toxic effector regulated by the C-terminus. Elife 6: pii: e23473. 10.7554/eLife.23473.</p>","pubmedId":"28527237","doi":""},{"reference":"<p>Frontzek K, Bardelli M, Senatore A, Henzi A, Reimann RR, Bedir S, et al., Aguzzi A. 2022. A conformational switch controlling the toxicity of the prion protein. Nat Struct Mol Biol 29(8): 831-840.</p>","pubmedId":"35948768","doi":""},{"reference":"<p>Mercer RCC, Harris DA. 2023. Mechanisms of prion-induced toxicity. Cell Tissue Res 392(1): 81-96.</p>","pubmedId":"36070155","doi":""}],"title":"<p>Against All Odds: Computational Screening Via Machine Learning Ranking and Generation of Antibody Candidates for Creutzfeldt-Jakob Disease</p>","reviews":[],"curatorReviews":[]},{"id":"eb15a4d2-6d02-4cab-8641-be9d3022aa71","decision":"revise","abstract":"<p>Creutzfeldt-Jakob disease (CJD) is a fatal prion disorder with no approved treatments. This study developed a machine learning pipeline trained on 21 literature-curated PrP-targeting CDR sequences to rank 29,574 original and mutation-generated candidate sequences using neutralization, selectivity, and literature-based blood-brain barrier (BBB) proxy scores. The models showed internal performance (neutralization AUC = 0.9239; selectivity AUC = 0.7759), identified 10 high-percentile candidates, and generated 7,309 novel variants. These findings support hypothesis-generating sequence-level prioritization of PrP-targeting candidates, but do not establish native PrP<sup>Sc</sup>-specific binding, full-antibody efficacy, exact PrP epitope recognition, or in vivo BBB penetration and require further experimental validation and testing.</p>","acknowledgements":"<p>The authors would like to thank Academy of the Canyons for providing resources to enable this project.</p>","authors":[{"affiliations":["College of the Canyons, Santa Clarita, California, United States","Academy of the Canyons, Santa Clarita, California, United States"],"departments":["",""],"credit":["conceptualization","dataCuration","formalAnalysis","fundingAcquisition","investigation","methodology","project","resources","software","validation","visualization","writing_originalDraft","writing_reviewEditing"],"email":"vtiruvellore@my.canyons.edu","firstName":"Vinayan","lastName":"Tiruvellore","submittingAuthor":true,"correspondingAuthor":true,"equalContribution":false,"WBId":null,"orcid":"https://orcid.org/0009-0003-5451-2503"},{"affiliations":["Academy of the Canyons, Santa Clarita, California, United States","College of the Canyons, Santa Clarita, California, United States"],"departments":["",""],"credit":["supervision","resources"],"email":"mkoegle@hartdistrict.org","firstName":"Mike","lastName":"Koegle","submittingAuthor":false,"correspondingAuthor":false,"equalContribution":false,"WBId":null,"orcid":null}],"awards":[],"conflictsOfInterest":"<p>The authors declare that there are no conflicts of interest present.</p>","dataTable":{"url":null},"extendedData":[],"funding":"<p>This work did not receive external funding. Support was provided by the authors and Academy of the Canyons, Santa Clarita, CA.</p>","image":{"url":"https://portal.micropublication.org/uploads/704d795891ef24bc74c9e5699dd42d7f.png"},"imageCaption":"<p>A) Machine learning pipeline for sequence-level computational prioritization of PrP-targeting CDR candidates. flow chart overview of the computational workflow used to evaluate CDR sequences against three screening criteria: blood-brain barrier compatibility proxy, selectivity-like scoring, and predicted neutralization-like activity. Background antibody sequences were scored relative to PrP-targeting exemplars, ranked by composite candidate-prioritization score, and used as the basis for mutation-based generation of novel candidate variants.</p><p>B) Feature engineering and sequence-level model inputs. Quantitative physicochemical descriptors were extracted from antibody CDR sequences, including length, hydropathy, residue composition, charge distribution, aromatic fraction, and glycine/proline content. These engineered sequence features were used to train classifiers that distinguish PrP-targeting exemplar sequences from non-PrP background sequences and support multi-objective ranking.</p><p>C) Model performance and candidate ranking results. Predicted neutralization-like and selectivity-like scores separated curated PrP-targeting positive exemplars from background antibodies and FDA-approved CNS antibody controls. ROC AUC and Mann-Whitney U test results support statistically significant discrimination between PrP-targeting exemplar sequences and non-PrP comparison groups.</p><p>D) Structural comparison of parent and mutation-derived antibody candidates. ESMFold-predicted CDR structures show conformational effects of single-residue deletions in high-scoring mutation-derived variants. Removal of structurally restrictive residues altered loop geometry and corresponded with improved composite computational ranking relative to the parent sequences.</p>","imageTitle":"<p>Machine Learning-Based Sequence Prioritization and Generation of Candidate Antibody CDR Sequences</p>","methods":"<p>Data collection: This study was designed as a preliminary computational screening pipeline for candidate PrP-targeting CDR sequences rather than as evidence of therapeutic efficacy. The positive set consisted of 21 curated PrP-targeting CDR sequences compiled from patents and literature-derived criteria sets: 8 sequences in the neutralization set, 11 in the selectivity set, and 2 BBB-related comparator sequences. The modeled sequences represent variable-length CDR segments ranging from 5 to 17 amino acids and therefore capture only CDR-derived subregions of the antibody variable domain rather than a complete antibody molecule. Because experimentally validated anti-PrP antibody datasets are scarce, this curated set was used as a small hypothesis-generating positive reference set. For model contrast, more than 20,000 non-PrP human IGH background CDR sequences were used as negatives during training. Additionally, the source reference, sequence identity, exact sequence used, region, and model role for each exemplar are provided in exemplar_table_full.csv in the project GitHub repository.</p><p>Control construction: The primary negative control was the non-PrP IGH background set. An external comparator set consisting of 8 FDA-approved CNS-related antibodies that were not designed to bind PrP was used during downstream evaluation. These FDA CNS antibodies were treated as non-PrP comparators to test whether the scoring framework was capturing PrP-related sequence patterns rather than general CNS-targeting characteristics.</p><p>Feature engineering: Each CDR sequence was converted into a hybrid feature representation consisting of character-level 3-mer sequence features and 7 physicochemical features. The physicochemical features were length, hydrophobic fraction, aromatic fraction, positive-residue fraction, negative-residue fraction, net charge per residue, and glycine/proline fraction. In the final target feature matrix, the 21 positive sequences were represented by 137 distinct 3-mer features and 7 physicochemical features. This representation was chosen to preserve short sequence motif information while also capturing coarse biochemical properties relevant to antibody recognition and developability.</p><p>Machine-learning approach: Two binary classifiers were trained: a neutralization-like classifier and a selectivity-like classifier. Both models used L2-regularized logistic regression with class balancing and the liblinear solver. Positive labels were assigned from the curated PrP-targeting sequence set, and negative labels were assigned from the non-PrP background set. The models therefore estimated whether a sequence more closely resembled the curated neutralization or selectivity exemplars than the background repertoire. The computational framework was trained on sequence-derived features from curated PrP-targeting CDR exemplars and was not trained directly against the full human PrP<sup>C</sup> sequence. It does not predict a specific PrP binding epitope. These outputs should therefore be interpreted as ranking scores rather than direct predictions of full-antibody therapeutic activity or native PrP<sup>Sc</sup>-specific binding.</p><p>Training and validation: Model performance was evaluated using 5-fold stratified cross-validation with ROC AUC and average precision. The neutralization-like classifier achieved mean ROC AUC = 0.9239 ± 0.0810, and the selectivity-like classifier achieved mean ROC AUC = 0.7759 ± 0.1217 across folds. Because only 21 curated positive exemplars were available, these models should be interpreted as preliminary ranking models rather than definitive predictive classifiers. The larger variance observed for the selectivity model is consistent with reduced stability under limited positive-sample conditions.</p><p>BBB proxy scoring model: The BBB proxy scoring model was a hand-engineered sequence-ranking function rather than an experimentally trained BBB penetration model. For each antibody sequence, the model computed six physicochemical descriptors: sequence length, approximate molecular weight, net charge at pH 7.4, estimated isoelectric point (pI), GRAVY hydropathy (Kyte-Doolittle average hydropathy), and a hydrophobic patch score defined as the fraction of 6-residue windows containing at least 4 hydrophobic residues. The proxy then assigned a BBB compatibility score by rewarding sequences with net charge near 6.0 and pI near 9.6 using Gaussian-like soft-band terms, while penalizing excessive charge, excessively high pI, large molecular weight, high hydropathy, and high hydrophobic patchiness using smooth soft-penalty terms. The raw score was computed as:</p><p>0.35 x charge_reward + 0.25 x pI_reward - 0.15 x size_penalty - 0.15 x gravy_penalty - 0.10 x patch_penalty - 0.10 x charge_extreme_penalty - 0.10 x pI_extreme_penalty</p><p>and then transformed to a 0-1 scale with the logistic function where BBB_score = 1 / (1 + e<sup>(-4 * (raw - 0.3)</sup>)). Higher scores indicated stronger predicted BBB compatibility under this proxy, whereas lower scores indicated weaker compatibility. Because no independent BBB transcytosis or transport training dataset was used, this score should be interpreted strictly as a computational ranking proxy for sequence-level BBB compatibility rather than as direct evidence that intact antibodies would penetrate the BBB in vivo.</p><p>Statistical adequacy and overfitting: The positive reference set was small, so overfitting risk is an important limitation. To mitigate this, the classifiers used regularized logistic regression, class balancing, and 5-fold stratified cross-validation, and were evaluated against both a large background repertoire and an external FDA CNS comparator set. These steps support the use of the framework for preliminary prioritization, but they do not eliminate the need for independent experimental validation. Accordingly, all outputs should be interpreted as hypothesis-generating candidate rankings rather than validated therapeutic predictions.</p><p>Candidate ranking: Candidate sequences were ranked by combining the neutralization-like score, the selectivity-like score, and the BBB proxy score. Ranking was performed on original seed sequences and mutation-generated variants. Across 29,574 evaluated sequences in the ranking/validation workflow, the top candidates were prioritized based on combined percentile performance across the three criteria.</p><p>Statistical testing: Mann-Whitney U tests were used to compare score distributions between predefined groups, including positives versus background and positives versus FDA CNS comparators, for both neutralization-like and selectivity-like scores. These tests were used as score-separation analyses rather than proof of causal biological discrimination or native PrP<sup>Sc</sup>-specific binding.</p><p>Mutation modeling: Sequence diversification was performed computationally by targeted mutation and rescoring. This procedure generated 7,309 novel sequences not present in the original library. These mutated candidates were evaluated using the same ranking framework as the original sequences.</p><p>Data and code availability: The curated input datasets, extracted feature tables, trained scoring artifacts, ranked candidate outputs, analysis code, and the exemplar table with the CDR sequences the model was trained on are publicly available at https://github.com/Axy-lis/CJD_ML_AAO.</p>","reagents":"<p></p>","patternDescription":"<p>CJD is a rare but universally fatal prion disease. There is no cure, no disease-modifying therapy. CJD is caused by the misfolding of native prion protein, PrP<sup>C</sup>, into the pathogenic isoform PrP<sup>Sc</sup>. PrP<sup>Sc</sup> acts as a self-propagating template, converts normal proteins into misfolded forms, forms amyloid aggregates, and causes widespread neuronal death (Prusiner 1998; Aguzzi and Calella 2009). Therapeutic development is challenging because the pathogenic PrP<sup>Sc</sup> fold can differ substantially between species and disease-associated prion strains, so compounds effective against mouse prions may fail against human prions. Additionally, the blood-brain barrier (BBB) limits CNS delivery, and therapies must avoid disrupting normal PrP<sup>C</sup> function. Despite decades of research, no clinically approved antibody or small-molecule therapy exists (Aguzzi and Calella 2009; Pankiewicz et al. 2019).</p><p>Many reported anti-PrP antibodies act through binding to PrP<sup>C</sup> or related conformational states, and apparent PrP<sup>Sc</sup> recognition can depend on denaturation or assay-specific exposure of epitopes. Accordingly, the present study should not be interpreted as establishing selective binding to native PrP<sup>Sc</sup>. In addition, prior studies have reported neurotoxicity associated with some anti-PrP antibodies and related PrP toxic signaling pathways, emphasizing that antibody-based intervention in prion disease remains biologically complex and that the present computational rankings should not be interpreted as evidence of safety or therapeutic suitability (Sonati et al. 2013; Reimann et al. 2016; Wu et al. 2017; Frontzek et al. 2022; Mercer and Harris 2023).</p><p>Machine learning approaches to antibody engineering are promising in identifying candidate antibodies. By training ranking models on sequence-derived CDR features from curated anti-PrP antibody exemplars, it is possible to evaluate large candidate libraries and identify those most likely to satisfy computational screening criteria. The experimental question was whether a machine learning-driven computational pipeline could prioritize candidate PrP-targeting CDR segments for CJD using neutralization-like, selectivity-like, and BBB compatibility proxy criteria. The hypothesis was that a machine learning pipeline trained on sequence-derived features from literature-curated PrP-targeting CDR exemplars would rank anti-PrP-like candidates above background sequences, distinguish them from unrelated CNS antibodies, and generate novel, high-scoring variants through mutation modeling. The modeled inputs in this study were variable-length CDR segments rather than complete antibody molecules or fixed-length 10-mers. In the curated exemplar set, these CDR sequences ranged from 5 to 17 amino acids, with many clustering near approximately 10 amino acids. The framework was therefore designed for sequence-level prioritization rather than full-antibody functional prediction.</p><p>The primary control was baseline background antibodies. An external negative control used FDA-approved CNS-targeting antibodies to test whether the model captured anti-PrP sequence signatures rather than general CNS-targeting characteristics. For data collection, 21 literature-curated PrP-targeting CDR exemplars were used as positives, and more than 20,000 non-PrP sequences were used as background. Anti-PrP antibody examples and PrP-targeting antibody concepts were drawn from prior literature and patent sources (Pankiewicz et al. 2019; Uger et al. 2016; Williamson et al. 2001; Collinge and Hawke 2009). Sequences derived from Uger et al. 2016 were included as literature-reported PrP-targeting exemplars for computational comparison, but not as definitive evidence of selective native PrP<sup>Sc</sup> recognition. For feature engineering, CDR sequences were converted into quantitative descriptors, including 3-mer motif embeddings, amino acid composition, length, net charge at pH 7.4, hydropathy using the Kyte-Doolittle scale, aromatic fraction, and residue class proportions.</p><p>Model training was designed to distinguish positives from background. The pipeline included a neutralization-like classifier, a selectivity-like classifier, and a BBB compatibility proxy. Candidate ranking was performed by scoring more than 20,000 sequences and calculating a composite score across all three criteria. The top 25 candidates were selected. Mutation modeling was then applied through targeted sequence modifications, and modified sequences were re-scored to identify novel high-ranking candidates.</p><p>The neutralization-like classifier achieved ROC AUC = 0.9239 ± 0.0810, and the selectivity-like classifier achieved ROC AUC = 0.7759 ± 0.1217. Both were above random expectation, supporting that CDR sequence-derived features carry meaningful signal for PrP-targeting exemplar ranking. The BBB compatibility proxy was a literature-based heuristic informed by papers on antibody BBB transport and anti-PrP antibody fragments (Triguero et al. 1989; Yang et al. 2014; Ruiz-López et al. 2021). Strong internal model performance was achieved despite training on only 21 curated exemplars.</p><p>Mann-Whitney U tests showed statistically significant score separation between positive exemplars and background sequences and between positives and FDA CNS antibodies. Positives versus background for neutralization had a U statistic of 180000 and a p-value of 0.000000102. Positives versus background for selectivity had a U statistic of 300000 and a p-value of 0.0000000001. Positives versus FDA CNS antibodies for neutralization had a U statistic of 72 and a p-value of 0.0003129. Positives versus FDA CNS antibodies for selectivity had a U statistic of 120 and a p-value of 0.0000609. These low p-values support score separation between predefined groups, but do not independently establish antibody activity, selective native PrP<sup>Sc</sup> recognition, or guarantee model generalizability.</p><p>The neutralization score distributions showed positive exemplars clustered near the highest predicted values, around 0.99 to 1.0, whereas background sequences remained spread across lower score values, around 0.3 to 0.7. The selectivity-like score distribution showed a similar pattern, with positive exemplars again occupying the highest scoring region and background sequences occupying lower score ranges. Neutralization-like and selectivity-like scores were also compared across known PrP-targeting exemplars, FDA CNS antibodies, and background sequences. Positive exemplars clustered near maximal scores, FDA CNS antibodies scored near zero, and background sequences showed low-score distributions. Together, these comparisons support that the classifiers captured features associated with the PrP-targeting reference set rather than general antibody features.</p><p>The ranked candidate pool consisted of more than 20,000 original antibody sequences and about 7,000 mutation-generated variants. This search space allowed the machine learning models to identify rare high-scoring candidates within diverse antibody sequence space. Ten candidates ranked in the 95th to 99th percentile across 29,574 evaluated sequences. Positive exemplars were heavily enriched in the top 50, while FDA CNS antibodies were entirely absent. Mutation modeling generated 7,309 novel sequences not present in the original library. The top novel candidates ranked 3rd, 7th, and 12th overall, outperforming 99.97% of all evaluated sequences in this computational screen.</p><p>Single-residue deletions produced the largest improvements. In one case, parent sequence QQNKNWPPGT ranked 17, while novel sequence QQNKNWPGT ranked 3. A single Pro8 deletion distinguished the parent from the novel generated sequence. Removing Pro8 relieved local conformational constraint in the CDR-segment loop and improved the composite computational score from 0.9738 to 0.9810, advancing the sequence from rank 17 to rank 3 out of 29,574 total candidates. In another case, parent sequence SSYTITNTQK ranked 367, while novel sequence SSTITNTQK ranked 12. A single Tyr3 deletion distinguished the parent from the novel generated sequence and increased the combined score from 0.9439 to 0.9759, producing the largest score improvement across the highlighted mutation pairs. These examples suggest that small local sequence changes can measurably alter the ranking outputs of the computational framework.</p><p>This study shows that machine learning can prioritize and refine candidate CDR sequences for Creutzfeldt-Jakob disease despite extremely limited training data. Both predictive classifiers distinguished PrP-targeting exemplars from unrelated background sequences using only sequence-derived features. The neutralization-like model achieved strong discriminative performance, while the selectivity-like classifier demonstrated moderate yet statistically significant predictive capability. These results indicate that meaningful sequence-level signal is encoded within antibody sequence features such as residue composition, charge distribution, hydrophobicity, and short motif content.</p><p>Composite ranking across three independent computational screening constraints, including BBB compatibility proxy, anti-PrP selectivity-like scoring, and neutralization potential, allowed the pipeline to approximate candidate-prioritization trade-offs (Dobson et al. 2016; Mieczkowski et al. 2023). Rather than optimizing a single metric, the framework balances multiple biological requirements simultaneously. This multi-objective scoring approach enabled identification of candidates that perform well across all criteria rather than excelling in only one dimension.</p><p>External control comparison using FDA-approved CNS antibodies provided additional evidence that the models capture features associated with PrP-targeting rather than generic CNS-associated properties. These control antibodies scored near zero in both predictive models, while PrP-targeting exemplars clustered near the highest score ranges. This separation supports that the classifiers learned biologically relevant signal associated with PrP targeting rather than simply recognizing general CNS antibody traits.</p><p>The mutation model further expanded the search space by generating targeted sequence modifications around top-ranking candidates. Several mutation-derived variants achieved scores comparable to or exceeding their parent sequences, demonstrating the potential for computational prioritization of candidate antibody sequences. Traditional anti-PrP drug discovery is bottlenecked by the scarcity of validated antibody datasets and the difficulty of experimental prion screening. This pipeline demonstrates that sequence-level machine learning can extract meaningful biological signal from as few as 21 positive examples, but the findings remain preliminary and computational. The identified candidates represent a prioritized set for later validation, including molecular docking against PrP structures, binding affinity assays, cell-based BBB transcytosis models, and prion cell model neutralization testing. Future work would center around gathering more positive exemplars, wet-lab validation, expanding the mutation model, and replacing the BBB proxy with a trained transport model. These results should be interpreted as hypothesis for sequence ranking rather than proof of native PrP<sup>Sc</sup>-specific binding, defined PrP epitope recognition, or BBB permeation. These results should also be interpreted in light of prior reports of anti-PrP antibody-associated neurotoxicity.</p>","references":[{"reference":"Aguzzi A, Calella AM. 2009. Prions: Protein aggregation and infectious diseases. Physiological Reviews. 89: 1105.","pubmedId":"","doi":"10.1152/physrev.00006.2009"},{"reference":"Dobson CL, Devine PW, Phillips JJ, Higazi DR, Lloyd C, Popplewell AG. 2016. Engineering the surface properties of a human monoclonal antibody prevents self-association and rapid clearance in vivo. Scientific Reports. 6: 38644.","pubmedId":"","doi":"10.1038/srep38644"},{"reference":"Mieczkowski C, Zhang X, Lee D, Nguyen K, Lv W, Wang Y, et al., Gries JM. 2023. Blueprint for antibody biologics developability. mAbs. 15: 2185924.","pubmedId":"","doi":"10.1080/19420862.2023.2185924"},{"reference":"Pankiewicz JE, Sanchez S, Kirshenbaum K, Kascsak RB, Kascsak RJ, Sadowski MJ. 2019. Anti-prion protein antibody 6D11 restores cellular proteostasis of prion protein through disrupting recycling propagation of PrPSc and targeting PrPSc for lysosomal degradation. Molecular Neurobiology. 56: 2073.","pubmedId":"","doi":"10.1007/s12035-018-1208-4"},{"reference":"Prusiner SB. 1998. Prions. Proceedings of the National Academy of Sciences. 95: 13363.","pubmedId":"","doi":"10.1073/pnas.95.23.13363"},{"reference":"<p>Ruiz-López E, Schuhmacher AJ. 2021. Transportation of Single-Domain Antibodies through the Blood-Brain Barrier. Biomolecules 11(8): 10.3390/biom11081131.</p>","pubmedId":"34439797","doi":""},{"reference":"Triguero D, Buciak JB, Pardridge WM. 1989. Blood-brain barrier transport of cationized immunoglobulin G. Proceedings of the National Academy of Sciences. 86: 4761.","pubmedId":"","doi":"10.1073/pnas.86.12.4761"},{"reference":"Uger MD, Chai V, Ciolfi V, Cashman NR, Tian B, Wong WY, Chao HLM. 2016. Antibodies and conjugates that target misfolded prion protein.","pubmedId":"","doi":""},{"reference":"Williamson RA, Burton DR, Prusiner SB. 2001. Antibodies specific for native PrPSc.","pubmedId":"","doi":""},{"reference":"Yang W. 2014. Efficient penetration of the blood-brain barrier by anti-prion antibody fragments. Science Translational Medicine. 6: 234ra60.","pubmedId":"","doi":"10.1126/scitranslmed.3008870"},{"reference":"Collinge J, Hawke S. 2009. Prion inhibition.","pubmedId":"","doi":""},{"reference":"<p>Sonati T, Reimann RR, Falsig J, Baral PK, O'Connor T, Hornemann S, et al., Aguzzi A. 2013. The toxicity of antiprion antibodies is mediated by the flexible tail of the prion protein. Nature 501(7465): 102-6.</p>","pubmedId":"23903654","doi":""},{"reference":"<p>Reimann RR, Sonati T, Hornemann S, Herrmann US, Arand M, Hawke S, Aguzzi A. 2016. Differential Toxicity of Antibodies to the Prion Protein. PLoS Pathog 12(1): e1005401.</p>","pubmedId":"26821311","doi":""},{"reference":"<p>Wu B, McDonald AJ, Markham K, Rich CB, McHugh KP, Tatzelt J, et al., Harris DA. 2017. The N-terminus of the prion protein is a toxic effector regulated by the C-terminus. Elife 6: pii: e23473. 10.7554/eLife.23473.</p>","pubmedId":"28527237","doi":""},{"reference":"<p>Frontzek K, Bardelli M, Senatore A, Henzi A, Reimann RR, Bedir S, et al., Aguzzi A. 2022. A conformational switch controlling the toxicity of the prion protein. Nat Struct Mol Biol 29(8): 831-840.</p>","pubmedId":"35948768","doi":""},{"reference":"<p>Mercer RCC, Harris DA. 2023. Mechanisms of prion-induced toxicity. Cell Tissue Res 392(1): 81-96.</p>","pubmedId":"36070155","doi":""}],"title":"<p>Against All Odds: Computational Screening Via Machine Learning Ranking and Generation of Antibody Candidates for Creutzfeldt-Jakob Disease</p>","reviews":[],"curatorReviews":[]},{"id":"b82bed54-c85c-4568-a040-38a5ec1bbfd4","decision":"edit","abstract":"<p>Creutzfeldt-Jakob disease (CJD) is a fatal prion disorder with no approved treatments. This study developed a machine learning pipeline trained on 21 literature-curated PrP-targeting CDR sequences to rank 29,574 original and mutation-generated candidate sequences using neutralization, selectivity, and literature-based blood-brain barrier (BBB) proxy scores. The models showed internal performance (neutralization AUC = 0.9239; selectivity AUC = 0.7759), identified 10 high-percentile candidates, and generated 7,309 novel variants. These findings support hypothesis-generating sequence-level prioritization of PrP-targeting candidates, but do not establish native PrP<sup>Sc</sup>-specific binding, full-antibody efficacy, exact PrP epitope recognition, or in vivo BBB penetration and require further experimental validation and testing.</p>","acknowledgements":"<p>The authors would like to thank Academy of the Canyons for providing resources to enable this project.</p>","authors":[{"affiliations":["College of the Canyons, Santa Clarita, California, United States","Academy of the Canyons, Santa Clarita, California, United States"],"departments":["",""],"credit":["conceptualization","dataCuration","formalAnalysis","fundingAcquisition","investigation","methodology","project","resources","software","validation","visualization","writing_originalDraft","writing_reviewEditing"],"email":"vtiruvellore@my.canyons.edu","firstName":"Vinayan","lastName":"Tiruvellore","submittingAuthor":true,"correspondingAuthor":true,"equalContribution":false,"WBId":null,"orcid":"https://orcid.org/0009-0003-5451-2503"},{"affiliations":["Academy of the Canyons, Santa Clarita, California, United States","College of the Canyons, Santa Clarita, California, United States"],"departments":["",""],"credit":["supervision","resources"],"email":"mkoegle@hartdistrict.org","firstName":"Mike","lastName":"Koegle","submittingAuthor":false,"correspondingAuthor":false,"equalContribution":false,"WBId":null,"orcid":null}],"awards":[],"conflictsOfInterest":"<p>The authors declare that there are no conflicts of interest present.</p>","dataTable":{"url":null},"extendedData":[],"funding":"<p>This work did not receive external funding. Support was provided by the authors and Academy of the Canyons, Santa Clarita, CA.</p>","image":{"url":"https://portal.micropublication.org/uploads/7a2ba7c6672cc12334c9fef3fb7eed6f.jpg"},"imageCaption":"<p>A) Machine learning pipeline for sequence-level computational prioritization of PrP-targeting CDR candidates. flow chart overview of the computational workflow used to evaluate CDR sequences against three screening criteria: blood-brain barrier compatibility proxy, selectivity-like scoring, and predicted neutralization-like activity. Background antibody sequences were scored relative to PrP-targeting exemplars, ranked by composite candidate-prioritization score, and used as the basis for mutation-based generation of novel candidate variants.</p><p>B) Feature engineering and sequence-level model inputs. Quantitative physicochemical descriptors were extracted from antibody CDR sequences, including length, hydropathy, residue composition, charge distribution, aromatic fraction, and glycine/proline content. These engineered sequence features were used to train classifiers that distinguish PrP-targeting exemplar sequences from non-PrP background sequences and support multi-objective ranking.</p><p>C) Model performance and candidate ranking results. Predicted neutralization-like and selectivity-like scores separated curated PrP-targeting positive exemplars from background antibodies and FDA-approved CNS antibody controls. ROC AUC and Mann-Whitney U test results support statistically significant discrimination between PrP-targeting exemplar sequences and non-PrP comparison groups.</p><p>D) Structural comparison of parent and mutation-derived antibody candidates. ESMFold-predicted CDR structures show conformational effects of single-residue deletions in high-scoring mutation-derived variants. Removal of structurally restrictive residues altered loop geometry and corresponded with improved composite computational ranking relative to the parent sequences.</p>","imageTitle":"<p>Machine Learning-Based Sequence Prioritization and Generation of Candidate Antibody CDR Sequences</p>","methods":"<p>Data collection: This study was designed as a preliminary computational screening pipeline for candidate PrP-targeting CDR sequences rather than as evidence of therapeutic efficacy. The positive set consisted of 21 curated PrP-targeting CDR sequences compiled from patents and literature-derived criteria sets: 8 sequences in the neutralization set, 11 in the selectivity set, and 2 BBB-related comparator sequences. The modeled sequences represent variable-length CDR segments ranging from 5 to 17 amino acids and therefore capture only CDR-derived subregions of the antibody variable domain rather than a complete antibody molecule. Because experimentally validated anti-PrP antibody datasets are scarce, this curated set was used as a small hypothesis-generating positive reference set. For model contrast, more than 20,000 non-PrP human IGH background CDR sequences were used as negatives during training. Additionally, the source reference, sequence identity, exact sequence used, region, and model role for each exemplar are provided in exemplar_table_full.csv in the project GitHub repository.</p><p>Control construction: The primary negative control was the non-PrP IGH background set. An external comparator set consisting of 8 FDA-approved CNS-related antibodies that were not designed to bind PrP was used during downstream evaluation. These FDA CNS antibodies were treated as non-PrP comparators to test whether the scoring framework was capturing PrP-related sequence patterns rather than general CNS-targeting characteristics.</p><p>Feature engineering: Each CDR sequence was converted into a hybrid feature representation consisting of character-level 3-mer sequence features and 7 physicochemical features. The physicochemical features were length, hydrophobic fraction, aromatic fraction, positive-residue fraction, negative-residue fraction, net charge per residue, and glycine/proline fraction. In the final target feature matrix, the 21 positive sequences were represented by 137 distinct 3-mer features and 7 physicochemical features. This representation was chosen to preserve short sequence motif information while also capturing coarse biochemical properties relevant to antibody recognition and developability.</p><p>Machine-learning approach: Two binary classifiers were trained: a neutralization-like classifier and a selectivity-like classifier. Both models used L2-regularized logistic regression with class balancing and the liblinear solver. Positive labels were assigned from the curated PrP-targeting sequence set, and negative labels were assigned from the non-PrP background set. The models therefore estimated whether a sequence more closely resembled the curated neutralization or selectivity exemplars than the background repertoire. The computational framework was trained on sequence-derived features from curated PrP-targeting CDR exemplars and was not trained directly against the full human PrP<sup>C</sup> sequence. It does not predict a specific PrP binding epitope. These outputs should therefore be interpreted as ranking scores rather than direct predictions of full-antibody therapeutic activity or native PrP<sup>Sc</sup>-specific binding.</p><p>Training and validation: Model performance was evaluated using 5-fold stratified cross-validation with ROC AUC and average precision. The neutralization-like classifier achieved mean ROC AUC = 0.9239 ± 0.0810, and the selectivity-like classifier achieved mean ROC AUC = 0.7759 ± 0.1217 across folds. Because only 21 curated positive exemplars were available, these models should be interpreted as preliminary ranking models rather than definitive predictive classifiers. The larger variance observed for the selectivity model is consistent with reduced stability under limited positive-sample conditions.</p><p>BBB proxy scoring model: The BBB proxy scoring model was a hand-engineered sequence-ranking function rather than an experimentally trained BBB penetration model. For each antibody sequence, the model computed six physicochemical descriptors: sequence length, approximate molecular weight, net charge at pH 7.4, estimated isoelectric point (pI), GRAVY hydropathy (Kyte-Doolittle average hydropathy), and a hydrophobic patch score defined as the fraction of 6-residue windows containing at least 4 hydrophobic residues. The proxy then assigned a BBB compatibility score by rewarding sequences with net charge near 6.0 and pI near 9.6 using Gaussian-like soft-band terms, while penalizing excessive charge, excessively high pI, large molecular weight, high hydropathy, and high hydrophobic patchiness using smooth soft-penalty terms. The raw score was computed as:</p><p>0.35 x charge_reward + 0.25 x pI_reward - 0.15 x size_penalty - 0.15 x gravy_penalty - 0.10 x patch_penalty - 0.10 x charge_extreme_penalty - 0.10 x pI_extreme_penalty</p><p>and then transformed to a 0-1 scale with the logistic function where BBB_score = 1 / (1 + e<sup>(-4 * (raw - 0.3)</sup>)). Higher scores indicated stronger predicted BBB compatibility under this proxy, whereas lower scores indicated weaker compatibility. Because no independent BBB transcytosis or transport training dataset was used, this score should be interpreted strictly as a computational ranking proxy for sequence-level BBB compatibility rather than as direct evidence that intact antibodies would penetrate the BBB in vivo.</p><p>Statistical adequacy and overfitting: The positive reference set was small, so overfitting risk is an important limitation. To mitigate this, the classifiers used regularized logistic regression, class balancing, and 5-fold stratified cross-validation, and were evaluated against both a large background repertoire and an external FDA CNS comparator set. These steps support the use of the framework for preliminary prioritization, but they do not eliminate the need for independent experimental validation. Accordingly, all outputs should be interpreted as hypothesis-generating candidate rankings rather than validated therapeutic predictions.</p><p>Candidate ranking: Candidate sequences were ranked by combining the neutralization-like score, the selectivity-like score, and the BBB proxy score. Ranking was performed on original seed sequences and mutation-generated variants. Across 29,574 evaluated sequences in the ranking/validation workflow, the top candidates were prioritized based on combined percentile performance across the three criteria.</p><p>Statistical testing: Mann-Whitney U tests were used to compare score distributions between predefined groups, including positives versus background and positives versus FDA CNS comparators, for both neutralization-like and selectivity-like scores. These tests were used as score-separation analyses rather than proof of causal biological discrimination or native PrP<sup>Sc</sup>-specific binding.</p><p>Mutation modeling: Sequence diversification was performed computationally by targeted mutation and rescoring. This procedure generated 7,309 novel sequences not present in the original library. These mutated candidates were evaluated using the same ranking framework as the original sequences.</p><p>Data and code availability: The curated input datasets, extracted feature tables, trained scoring artifacts, ranked candidate outputs, analysis code, and the exemplar table with the CDR sequences the model was trained on are publicly available at https://github.com/Axy-lis/CJD_ML_AAO.</p>","reagents":"<p></p>","patternDescription":"<p>CJD is a rare but universally fatal prion disease. There is no cure, no disease-modifying therapy. CJD is caused by the misfolding of native prion protein, PrP<sup>C</sup>, into the pathogenic isoform PrP<sup>Sc</sup>. PrP<sup>Sc</sup> acts as a self-propagating template, converts normal proteins into misfolded forms, forms amyloid aggregates, and causes widespread neuronal death (Prusiner 1998; Aguzzi and Calella 2009). Therapeutic development is challenging because the pathogenic PrP<sup>Sc</sup> fold can differ substantially between species and disease-associated prion strains, so compounds effective against mouse prions may fail against human prions. Additionally, the blood-brain barrier (BBB) limits CNS delivery, and therapies must avoid disrupting normal PrP<sup>C</sup> function. Despite decades of research, no clinically approved antibody or small-molecule therapy exists (Aguzzi and Calella 2009; Pankiewicz et al. 2019).</p><p>Many reported anti-PrP antibodies act through binding to PrP<sup>C</sup> or related conformational states, and apparent PrP<sup>Sc</sup> recognition can depend on denaturation or assay-specific exposure of epitopes. Accordingly, the present study should not be interpreted as establishing selective binding to native PrP<sup>Sc</sup>. In addition, prior studies have reported neurotoxicity associated with some anti-PrP antibodies and related PrP toxic signaling pathways, emphasizing that antibody-based intervention in prion disease remains biologically complex and that the present computational rankings should not be interpreted as evidence of safety or therapeutic suitability (Sonati et al. 2013; Reimann et al. 2016; Wu et al. 2017; Frontzek et al. 2022; Mercer and Harris 2023).</p><p>Machine learning approaches to antibody engineering are promising in identifying candidate antibodies. By training ranking models on sequence-derived CDR features from curated anti-PrP antibody exemplars, it is possible to evaluate large candidate libraries and identify those most likely to satisfy computational screening criteria. The experimental question was whether a machine learning-driven computational pipeline could prioritize candidate PrP-targeting CDR segments for CJD using neutralization-like, selectivity-like, and BBB compatibility proxy criteria. The hypothesis was that a machine learning pipeline trained on sequence-derived features from literature-curated PrP-targeting CDR exemplars would rank anti-PrP-like candidates above background sequences, distinguish them from unrelated CNS antibodies, and generate novel, high-scoring variants through mutation modeling. The modeled inputs in this study were variable-length CDR segments rather than complete antibody molecules or fixed-length 10-mers. In the curated exemplar set, these CDR sequences ranged from 5 to 17 amino acids, with many clustering near approximately 10 amino acids. The framework was therefore designed for sequence-level prioritization rather than full-antibody functional prediction.</p><p>The primary control was baseline background antibodies. An external negative control used FDA-approved CNS-targeting antibodies to test whether the model captured anti-PrP sequence signatures rather than general CNS-targeting characteristics. For data collection, 21 literature-curated PrP-targeting CDR exemplars were used as positives, and more than 20,000 non-PrP sequences were used as background. Anti-PrP antibody examples and PrP-targeting antibody concepts were drawn from prior literature and patent sources (Pankiewicz et al. 2019; Uger et al. 2016; Williamson et al. 2001; Collinge and Hawke 2009). Sequences derived from Uger et al. 2016 were included as literature-reported PrP-targeting exemplars for computational comparison, but not as definitive evidence of selective native PrP<sup>Sc</sup> recognition. For feature engineering, CDR sequences were converted into quantitative descriptors, including 3-mer motif embeddings, amino acid composition, length, net charge at pH 7.4, hydropathy using the Kyte-Doolittle scale, aromatic fraction, and residue class proportions.</p><p>Model training was designed to distinguish positives from background. The pipeline included a neutralization-like classifier, a selectivity-like classifier, and a BBB compatibility proxy. Candidate ranking was performed by scoring more than 20,000 sequences and calculating a composite score across all three criteria. The top 25 candidates were selected. Mutation modeling was then applied through targeted sequence modifications, and modified sequences were re-scored to identify novel high-ranking candidates.</p><p>The neutralization-like classifier achieved ROC AUC = 0.9239 ± 0.0810, and the selectivity-like classifier achieved ROC AUC = 0.7759 ± 0.1217. Both were above random expectation, supporting that CDR sequence-derived features carry meaningful signal for PrP-targeting exemplar ranking. The BBB compatibility proxy was a literature-based heuristic informed by papers on antibody BBB transport and anti-PrP antibody fragments (Triguero et al. 1989; Yang et al. 2014; Ruiz-López et al. 2021). Strong internal model performance was achieved despite training on only 21 curated exemplars.</p><p>Mann-Whitney U tests showed statistically significant score separation between positive exemplars and background sequences and between positives and FDA CNS antibodies. Positives versus background for neutralization had a U statistic of 180000 and a p-value of 0.000000102. Positives versus background for selectivity had a U statistic of 300000 and a p-value of 0.0000000001. Positives versus FDA CNS antibodies for neutralization had a U statistic of 72 and a p-value of 0.0003129. Positives versus FDA CNS antibodies for selectivity had a U statistic of 120 and a p-value of 0.0000609. These low p-values support score separation between predefined groups, but do not independently establish antibody activity, selective native PrP<sup>Sc</sup> recognition, or guarantee model generalizability.</p><p>The neutralization score distributions showed positive exemplars clustered near the highest predicted values, around 0.99 to 1.0, whereas background sequences remained spread across lower score values, around 0.3 to 0.7. The selectivity-like score distribution showed a similar pattern, with positive exemplars again occupying the highest scoring region and background sequences occupying lower score ranges. Neutralization-like and selectivity-like scores were also compared across known PrP-targeting exemplars, FDA CNS antibodies, and background sequences. Positive exemplars clustered near maximal scores, FDA CNS antibodies scored near zero, and background sequences showed low-score distributions. Together, these comparisons support that the classifiers captured features associated with the PrP-targeting reference set rather than general antibody features.</p><p>The ranked candidate pool consisted of more than 20,000 original antibody sequences and about 7,000 mutation-generated variants. This search space allowed the machine learning models to identify rare high-scoring candidates within diverse antibody sequence space. Ten candidates ranked in the 95th to 99th percentile across 29,574 evaluated sequences. Positive exemplars were heavily enriched in the top 50, while FDA CNS antibodies were entirely absent. Mutation modeling generated 7,309 novel sequences not present in the original library. The top novel candidates ranked 3rd, 7th, and 12th overall, outperforming 99.97% of all evaluated sequences in this computational screen.</p><p>Single-residue deletions produced the largest improvements. In one case, parent sequence QQNKNWPPGT ranked 17, while novel sequence QQNKNWPGT ranked 3. A single Pro8 deletion distinguished the parent from the novel generated sequence. Removing Pro8 relieved local conformational constraint in the CDR-segment loop and improved the composite computational score from 0.9738 to 0.9810, advancing the sequence from rank 17 to rank 3 out of 29,574 total candidates. In another case, parent sequence SSYTITNTQK ranked 367, while novel sequence SSTITNTQK ranked 12. A single Tyr3 deletion distinguished the parent from the novel generated sequence and increased the combined score from 0.9439 to 0.9759, producing the largest score improvement across the highlighted mutation pairs. These examples suggest that small local sequence changes can measurably alter the ranking outputs of the computational framework.</p><p>This study shows that machine learning can prioritize and refine candidate CDR sequences for Creutzfeldt-Jakob disease despite extremely limited training data. Both predictive classifiers distinguished PrP-targeting exemplars from unrelated background sequences using only sequence-derived features. The neutralization-like model achieved strong discriminative performance, while the selectivity-like classifier demonstrated moderate yet statistically significant predictive capability. These results indicate that meaningful sequence-level signal is encoded within antibody sequence features such as residue composition, charge distribution, hydrophobicity, and short motif content.</p><p>Composite ranking across three independent computational screening constraints, including BBB compatibility proxy, anti-PrP selectivity-like scoring, and neutralization potential, allowed the pipeline to approximate candidate-prioritization trade-offs (Dobson et al. 2016; Mieczkowski et al. 2023). Rather than optimizing a single metric, the framework balances multiple biological requirements simultaneously. This multi-objective scoring approach enabled identification of candidates that perform well across all criteria rather than excelling in only one dimension.</p><p>External control comparison using FDA-approved CNS antibodies provided additional evidence that the models capture features associated with PrP-targeting rather than generic CNS-associated properties. These control antibodies scored near zero in both predictive models, while PrP-targeting exemplars clustered near the highest score ranges. This separation supports that the classifiers learned biologically relevant signal associated with PrP targeting rather than simply recognizing general CNS antibody traits.</p><p>The mutation model further expanded the search space by generating targeted sequence modifications around top-ranking candidates. Several mutation-derived variants achieved scores comparable to or exceeding their parent sequences, demonstrating the potential for computational prioritization of candidate antibody sequences. Traditional anti-PrP drug discovery is bottlenecked by the scarcity of validated antibody datasets and the difficulty of experimental prion screening. This pipeline demonstrates that sequence-level machine learning can extract meaningful biological signal from as few as 21 positive examples, but the findings remain preliminary and computational. The identified candidates represent a prioritized set for later validation, including molecular docking against PrP structures, binding affinity assays, cell-based BBB transcytosis models, and prion cell model neutralization testing. Future work would center around gathering more positive exemplars, wet-lab validation, expanding the mutation model, and replacing the BBB proxy with a trained transport model. These results should be interpreted as hypothesis for sequence ranking rather than proof of native PrP<sup>Sc</sup>-specific binding, defined PrP epitope recognition, or BBB permeation. These results should also be interpreted in light of prior reports of anti-PrP antibody-associated neurotoxicity.</p>","references":[{"reference":"Aguzzi A, Calella AM. 2009. Prions: Protein aggregation and infectious diseases. Physiological Reviews. 89: 1105.","pubmedId":"","doi":"10.1152/physrev.00006.2009"},{"reference":"Dobson CL, Devine PW, Phillips JJ, Higazi DR, Lloyd C, Popplewell AG. 2016. Engineering the surface properties of a human monoclonal antibody prevents self-association and rapid clearance in vivo. Scientific Reports. 6: 38644.","pubmedId":"","doi":"10.1038/srep38644"},{"reference":"Mieczkowski C, Zhang X, Lee D, Nguyen K, Lv W, Wang Y, et al., Gries JM. 2023. Blueprint for antibody biologics developability. mAbs. 15: 2185924.","pubmedId":"","doi":"10.1080/19420862.2023.2185924"},{"reference":"Pankiewicz JE, Sanchez S, Kirshenbaum K, Kascsak RB, Kascsak RJ, Sadowski MJ. 2019. Anti-prion protein antibody 6D11 restores cellular proteostasis of prion protein through disrupting recycling propagation of PrPSc and targeting PrPSc for lysosomal degradation. Molecular Neurobiology. 56: 2073.","pubmedId":"","doi":"10.1007/s12035-018-1208-4"},{"reference":"Prusiner SB. 1998. Prions. Proceedings of the National Academy of Sciences. 95: 13363.","pubmedId":"","doi":"10.1073/pnas.95.23.13363"},{"reference":"<p>Ruiz-López E, Schuhmacher AJ. 2021. Transportation of Single-Domain Antibodies through the Blood-Brain Barrier. Biomolecules 11(8): 10.3390/biom11081131.</p>","pubmedId":"34439797","doi":""},{"reference":"Triguero D, Buciak JB, Pardridge WM. 1989. Blood-brain barrier transport of cationized immunoglobulin G. Proceedings of the National Academy of Sciences. 86: 4761.","pubmedId":"","doi":"10.1073/pnas.86.12.4761"},{"reference":"Uger MD, Chai V, Ciolfi V, Cashman NR, Tian B, Wong WY, Chao HLM. 2016. Antibodies and conjugates that target misfolded prion protein.","pubmedId":"","doi":""},{"reference":"Williamson RA, Burton DR, Prusiner SB. 2001. Antibodies specific for native PrPSc.","pubmedId":"","doi":""},{"reference":"Yang W. 2014. Efficient penetration of the blood-brain barrier by anti-prion antibody fragments. Science Translational Medicine. 6: 234ra60.","pubmedId":"","doi":"10.1126/scitranslmed.3008870"},{"reference":"Collinge J, Hawke S. 2009. Prion inhibition.","pubmedId":"","doi":""},{"reference":"<p>Sonati T, Reimann RR, Falsig J, Baral PK, O'Connor T, Hornemann S, et al., Aguzzi A. 2013. The toxicity of antiprion antibodies is mediated by the flexible tail of the prion protein. Nature 501(7465): 102-6.</p>","pubmedId":"23903654","doi":""},{"reference":"<p>Reimann RR, Sonati T, Hornemann S, Herrmann US, Arand M, Hawke S, Aguzzi A. 2016. Differential Toxicity of Antibodies to the Prion Protein. PLoS Pathog 12(1): e1005401.</p>","pubmedId":"26821311","doi":""},{"reference":"<p>Wu B, McDonald AJ, Markham K, Rich CB, McHugh KP, Tatzelt J, et al., Harris DA. 2017. The N-terminus of the prion protein is a toxic effector regulated by the C-terminus. Elife 6: pii: e23473. 10.7554/eLife.23473.</p>","pubmedId":"28527237","doi":""},{"reference":"<p>Frontzek K, Bardelli M, Senatore A, Henzi A, Reimann RR, Bedir S, et al., Aguzzi A. 2022. A conformational switch controlling the toxicity of the prion protein. Nat Struct Mol Biol 29(8): 831-840.</p>","pubmedId":"35948768","doi":""},{"reference":"<p>Mercer RCC, Harris DA. 2023. Mechanisms of prion-induced toxicity. Cell Tissue Res 392(1): 81-96.</p>","pubmedId":"36070155","doi":""}],"title":"<p>Against All Odds: Computational Screening Via Machine Learning Ranking and Generation of Antibody Candidates for Creutzfeldt-Jakob Disease</p>","reviews":[],"curatorReviews":[]},{"id":"e9b10642-8e57-4a4f-937d-e2e120768e85","decision":"edit","abstract":"<p>Creutzfeldt-Jakob disease (CJD) is a fatal prion disorder with no approved treatments. This study developed a machine learning pipeline trained on 21 literature-curated PrP-targeting CDR sequences to rank 29,574 original and mutation-generated candidate sequences using neutralization, selectivity, and literature-based blood-brain barrier (BBB) proxy scores. The models showed internal performance (neutralization AUC = 0.9239; selectivity AUC = 0.7759), identified 10 high-percentile candidates, and generated 7,309 novel variants. These findings support hypothesis-generating sequence-level prioritization of PrP-targeting candidates, but do not establish native PrP<sup>Sc</sup>-specific binding, full-antibody efficacy, exact PrP epitope recognition, or in vivo BBB penetration and require further experimental validation and testing.</p>","acknowledgements":"<p>The authors would like to thank Academy of the Canyons for providing resources to enable this project.</p>","authors":[{"affiliations":["College of the Canyons, Santa Clarita, California, United States","Academy of the Canyons, Santa Clarita, California, United States"],"departments":["",""],"credit":["conceptualization","dataCuration","formalAnalysis","fundingAcquisition","investigation","methodology","project","resources","software","validation","visualization","writing_originalDraft","writing_reviewEditing"],"email":"vtiruvellore@my.canyons.edu","firstName":"Vinayan","lastName":"Tiruvellore","submittingAuthor":true,"correspondingAuthor":true,"equalContribution":false,"WBId":null,"orcid":"https://orcid.org/0009-0003-5451-2503"},{"affiliations":["Academy of the Canyons, Santa Clarita, California, United States","College of the Canyons, Santa Clarita, California, United States"],"departments":["",""],"credit":["supervision","resources"],"email":"mkoegle@hartdistrict.org","firstName":"Mike","lastName":"Koegle","submittingAuthor":false,"correspondingAuthor":false,"equalContribution":false,"WBId":null,"orcid":null}],"awards":[],"conflictsOfInterest":"<p>The authors declare that there are no conflicts of interest present.</p>","dataTable":{"url":null},"extendedData":[],"funding":"<p>This work did not receive external funding. Support was provided by the authors and Academy of the Canyons, Santa Clarita, CA.</p>","image":{"url":"https://portal.micropublication.org/uploads/7a2ba7c6672cc12334c9fef3fb7eed6f.jpg"},"imageCaption":"<p>A) Machine learning pipeline for sequence-level computational prioritization of PrP-targeting CDR candidates. flow chart overview of the computational workflow used to evaluate CDR sequences against three screening criteria: blood-brain barrier compatibility proxy, selectivity-like scoring, and predicted neutralization-like activity. Background antibody sequences were scored relative to PrP-targeting exemplars, ranked by composite candidate-prioritization score, and used as the basis for mutation-based generation of novel candidate variants.</p><p>B) Feature engineering and sequence-level model inputs. Quantitative physicochemical descriptors were extracted from antibody CDR sequences, including length, hydropathy, residue composition, charge distribution, aromatic fraction, and glycine/proline content. These engineered sequence features were used to train classifiers that distinguish PrP-targeting exemplar sequences from non-PrP background sequences and support multi-objective ranking.</p><p>C) Model performance and candidate ranking results. Predicted neutralization-like and selectivity-like scores separated curated PrP-targeting positive exemplars from background antibodies and FDA-approved CNS antibody controls. ROC AUC and Mann-Whitney U test results support statistically significant discrimination between PrP-targeting exemplar sequences and non-PrP comparison groups.</p><p>D) Structural comparison of parent and mutation-derived antibody candidates. ESMFold-predicted CDR structures show conformational effects of single-residue deletions in high-scoring mutation-derived variants. Removal of structurally restrictive residues altered loop geometry and corresponded with improved composite computational ranking relative to the parent sequences.</p>","imageTitle":"<p>Machine Learning-Based Sequence Prioritization and Generation of Candidate Antibody CDR Sequences</p>","methods":"<p>Data collection: This study was designed as a preliminary computational screening pipeline for candidate PrP-targeting CDR sequences rather than as evidence of therapeutic efficacy. The positive set consisted of 21 curated PrP-targeting CDR sequences compiled from patents and literature-derived criteria sets: 8 sequences in the neutralization set, 11 in the selectivity set, and 2 BBB-related comparator sequences. The modeled sequences represent variable-length CDR segments ranging from 5 to 17 amino acids and therefore capture only CDR-derived subregions of the antibody variable domain rather than a complete antibody molecule. Because experimentally validated anti-PrP antibody datasets are scarce, this curated set was used as a small hypothesis-generating positive reference set. For model contrast, more than 20,000 non-PrP human IGH background CDR sequences were used as negatives during training. Additionally, the source reference, sequence identity, exact sequence used, region, and model role for each exemplar are provided in exemplar_table_full.csv in the project GitHub repository.</p><p>Control construction: The primary negative control was the non-PrP IGH background set. An external comparator set consisting of 8 FDA-approved CNS-related antibodies that were not designed to bind PrP was used during downstream evaluation. These FDA CNS antibodies were treated as non-PrP comparators to test whether the scoring framework was capturing PrP-related sequence patterns rather than general CNS-targeting characteristics.</p><p>Feature engineering: Each CDR sequence was converted into a hybrid feature representation consisting of character-level 3-mer sequence features and 7 physicochemical features. The physicochemical features were length, hydrophobic fraction, aromatic fraction, positive-residue fraction, negative-residue fraction, net charge per residue, and glycine/proline fraction. In the final target feature matrix, the 21 positive sequences were represented by 137 distinct 3-mer features and 7 physicochemical features. This representation was chosen to preserve short sequence motif information while also capturing coarse biochemical properties relevant to antibody recognition and developability.</p><p>Machine-learning approach: Two binary classifiers were trained: a neutralization-like classifier and a selectivity-like classifier. Both models used L2-regularized logistic regression with class balancing and the liblinear solver. Positive labels were assigned from the curated PrP-targeting sequence set, and negative labels were assigned from the non-PrP background set. The models therefore estimated whether a sequence more closely resembled the curated neutralization or selectivity exemplars than the background repertoire. The computational framework was trained on sequence-derived features from curated PrP-targeting CDR exemplars and was not trained directly against the full human PrP<sup>C</sup> sequence. It does not predict a specific PrP binding epitope. These outputs should therefore be interpreted as ranking scores rather than direct predictions of full-antibody therapeutic activity or native PrP<sup>Sc</sup>-specific binding.</p><p>Training and validation: Model performance was evaluated using 5-fold stratified cross-validation with ROC AUC and average precision. The neutralization-like classifier achieved mean ROC AUC = 0.9239 ± 0.0810, and the selectivity-like classifier achieved mean ROC AUC = 0.7759 ± 0.1217 across folds. Because only 21 curated positive exemplars were available, these models should be interpreted as preliminary ranking models rather than definitive predictive classifiers. The larger variance observed for the selectivity model is consistent with reduced stability under limited positive-sample conditions.</p><p>BBB proxy scoring model: The BBB proxy scoring model was a hand-engineered sequence-ranking function rather than an experimentally trained BBB penetration model. For each antibody sequence, the model computed six physicochemical descriptors: sequence length, approximate molecular weight, net charge at pH 7.4, estimated isoelectric point (pI), GRAVY hydropathy (Kyte-Doolittle average hydropathy), and a hydrophobic patch score defined as the fraction of 6-residue windows containing at least 4 hydrophobic residues. The proxy then assigned a BBB compatibility score by rewarding sequences with net charge near 6.0 and pI near 9.6 using Gaussian-like soft-band terms, while penalizing excessive charge, excessively high pI, large molecular weight, high hydropathy, and high hydrophobic patchiness using smooth soft-penalty terms. The raw score was computed as:</p><p>0.35 x charge_reward + 0.25 x pI_reward - 0.15 x size_penalty - 0.15 x gravy_penalty - 0.10 x patch_penalty - 0.10 x charge_extreme_penalty - 0.10 x pI_extreme_penalty</p><p>and then transformed to a 0-1 scale with the logistic function where BBB_score = 1 / (1 + e<sup>(-4 * (raw - 0.3)</sup>)). Higher scores indicated stronger predicted BBB compatibility under this proxy, whereas lower scores indicated weaker compatibility. Because no independent BBB transcytosis or transport training dataset was used, this score should be interpreted strictly as a computational ranking proxy for sequence-level BBB compatibility rather than as direct evidence that intact antibodies would penetrate the BBB in vivo.</p><p>Statistical adequacy and overfitting: The positive reference set was small, so overfitting risk is an important limitation. To mitigate this, the classifiers used regularized logistic regression, class balancing, and 5-fold stratified cross-validation, and were evaluated against both a large background repertoire and an external FDA CNS comparator set. These steps support the use of the framework for preliminary prioritization, but they do not eliminate the need for independent experimental validation. Accordingly, all outputs should be interpreted as hypothesis-generating candidate rankings rather than validated therapeutic predictions.</p><p>Candidate ranking: Candidate sequences were ranked by combining the neutralization-like score, the selectivity-like score, and the BBB proxy score. Ranking was performed on original seed sequences and mutation-generated variants. Across 29,574 evaluated sequences in the ranking/validation workflow, the top candidates were prioritized based on combined percentile performance across the three criteria.</p><p>Statistical testing: Mann-Whitney U tests were used to compare score distributions between predefined groups, including positives versus background and positives versus FDA CNS comparators, for both neutralization-like and selectivity-like scores. These tests were used as score-separation analyses rather than proof of causal biological discrimination or native PrP<sup>Sc</sup>-specific binding.</p><p>Mutation modeling: Sequence diversification was performed computationally by targeted mutation and rescoring. This procedure generated 7,309 novel sequences not present in the original library. These mutated candidates were evaluated using the same ranking framework as the original sequences.</p><p>Data and code availability: The curated input datasets, extracted feature tables, trained scoring artifacts, ranked candidate outputs, analysis code, and the exemplar table with the CDR sequences the model was trained on are publicly available at https://github.com/Axy-lis/CJD_ML_AAO.</p>","reagents":"<p></p>","patternDescription":"<p>CJD is a rare but universally fatal prion disease. There is no cure, no disease-modifying therapy. CJD is caused by the misfolding of native prion protein, PrP<sup>C</sup>, into the pathogenic isoform PrP<sup>Sc</sup>. PrP<sup>Sc</sup> acts as a self-propagating template, converts normal proteins into misfolded forms, forms amyloid aggregates, and causes widespread neuronal death (Prusiner 1998; Aguzzi and Calella 2009). Therapeutic development is challenging because the pathogenic PrP<sup>Sc</sup> fold can differ substantially between species and disease-associated prion strains, so compounds effective against mouse prions may fail against human prions. Additionally, the blood-brain barrier (BBB) limits CNS delivery, and therapies must avoid disrupting normal PrP<sup>C</sup> function. Despite decades of research, no clinically approved antibody or small-molecule therapy exists (Aguzzi and Calella 2009; Pankiewicz et al. 2019).</p><p>Many reported anti-PrP antibodies act through binding to PrP<sup>C</sup> or related conformational states, and apparent PrP<sup>Sc</sup> recognition can depend on denaturation or assay-specific exposure of epitopes. Accordingly, the present study should not be interpreted as establishing selective binding to native PrP<sup>Sc</sup>. In addition, prior studies have reported neurotoxicity associated with some anti-PrP antibodies and related PrP toxic signaling pathways, emphasizing that antibody-based intervention in prion disease remains biologically complex and that the present computational rankings should not be interpreted as evidence of safety or therapeutic suitability (Sonati et al. 2013; Reimann et al. 2016; Wu et al. 2017; Frontzek et al. 2022; Mercer and Harris 2023).</p><p>Machine learning approaches to antibody engineering are promising in identifying candidate antibodies. By training ranking models on sequence-derived CDR features from curated anti-PrP antibody exemplars, it is possible to evaluate large candidate libraries and identify those most likely to satisfy computational screening criteria. The experimental question was whether a machine learning-driven computational pipeline could prioritize candidate PrP-targeting CDR segments for CJD using neutralization-like, selectivity-like, and BBB compatibility proxy criteria. The hypothesis was that a machine learning pipeline trained on sequence-derived features from literature-curated PrP-targeting CDR exemplars would rank anti-PrP-like candidates above background sequences, distinguish them from unrelated CNS antibodies, and generate novel, high-scoring variants through mutation modeling. The modeled inputs in this study were variable-length CDR segments rather than complete antibody molecules or fixed-length 10-mers. In the curated exemplar set, these CDR sequences ranged from 5 to 17 amino acids, with many clustering near approximately 10 amino acids. The framework was therefore designed for sequence-level prioritization rather than full-antibody functional prediction.</p><p>The primary control was baseline background antibodies. An external negative control used FDA-approved CNS-targeting antibodies to test whether the model captured anti-PrP sequence signatures rather than general CNS-targeting characteristics. For data collection, 21 literature-curated PrP-targeting CDR exemplars were used as positives, and more than 20,000 non-PrP sequences were used as background. Anti-PrP antibody examples and PrP-targeting antibody concepts were drawn from prior literature and patent sources (Pankiewicz et al. 2019; Uger et al. 2016; Williamson et al. 2001; Collinge and Hawke 2009). Sequences derived from Uger et al. 2016 were included as literature-reported PrP-targeting exemplars for computational comparison, but not as definitive evidence of selective native PrP<sup>Sc</sup> recognition. For feature engineering, CDR sequences were converted into quantitative descriptors, including 3-mer motif embeddings, amino acid composition, length, net charge at pH 7.4, hydropathy using the Kyte-Doolittle scale, aromatic fraction, and residue class proportions.</p><p>Model training was designed to distinguish positives from background. The pipeline included a neutralization-like classifier, a selectivity-like classifier, and a BBB compatibility proxy. Candidate ranking was performed by scoring more than 20,000 sequences and calculating a composite score across all three criteria. The top 25 candidates were selected. Mutation modeling was then applied through targeted sequence modifications, and modified sequences were re-scored to identify novel high-ranking candidates.</p><p>The neutralization-like classifier achieved ROC AUC = 0.9239 ± 0.0810, and the selectivity-like classifier achieved ROC AUC = 0.7759 ± 0.1217. Both were above random expectation, supporting that CDR sequence-derived features carry meaningful signal for PrP-targeting exemplar ranking. The BBB compatibility proxy was a literature-based heuristic informed by papers on antibody BBB transport and anti-PrP antibody fragments (Triguero et al. 1989; Yang et al. 2014; Ruiz-López et al. 2021). Strong internal model performance was achieved despite training on only 21 curated exemplars.</p><p>Mann-Whitney U tests showed statistically significant score separation between positive exemplars and background sequences and between positives and FDA CNS antibodies. Positives versus background for neutralization had a U statistic of 180000 and a p-value of 0.000000102. Positives versus background for selectivity had a U statistic of 300000 and a p-value of 0.0000000001. Positives versus FDA CNS antibodies for neutralization had a U statistic of 72 and a p-value of 0.0003129. Positives versus FDA CNS antibodies for selectivity had a U statistic of 120 and a p-value of 0.0000609. These low p-values support score separation between predefined groups, but do not independently establish antibody activity, selective native PrP<sup>Sc</sup> recognition, or guarantee model generalizability.</p><p>The neutralization score distributions showed positive exemplars clustered near the highest predicted values, around 0.99 to 1.0, whereas background sequences remained spread across lower score values, around 0.3 to 0.7. The selectivity-like score distribution showed a similar pattern, with positive exemplars again occupying the highest scoring region and background sequences occupying lower score ranges. Neutralization-like and selectivity-like scores were also compared across known PrP-targeting exemplars, FDA CNS antibodies, and background sequences. Positive exemplars clustered near maximal scores, FDA CNS antibodies scored near zero, and background sequences showed low-score distributions. Together, these comparisons support that the classifiers captured features associated with the PrP-targeting reference set rather than general antibody features.</p><p>The ranked candidate pool consisted of more than 20,000 original antibody sequences and about 7,000 mutation-generated variants. This search space allowed the machine learning models to identify rare high-scoring candidates within diverse antibody sequence space. Ten candidates ranked in the 95th to 99th percentile across 29,574 evaluated sequences. Positive exemplars were heavily enriched in the top 50, while FDA CNS antibodies were entirely absent. Mutation modeling generated 7,309 novel sequences not present in the original library. The top novel candidates ranked 3rd, 7th, and 12th overall, outperforming 99.97% of all evaluated sequences in this computational screen.</p><p>Single-residue deletions produced the largest improvements. In one case, parent sequence QQNKNWPPGT ranked 17, while novel sequence QQNKNWPGT ranked 3. A single Pro8 deletion distinguished the parent from the novel generated sequence. Removing Pro8 relieved local conformational constraint in the CDR-segment loop and improved the composite computational score from 0.9738 to 0.9810, advancing the sequence from rank 17 to rank 3 out of 29,574 total candidates. In another case, parent sequence SSYTITNTQK ranked 367, while novel sequence SSTITNTQK ranked 12. A single Tyr3 deletion distinguished the parent from the novel generated sequence and increased the combined score from 0.9439 to 0.9759, producing the largest score improvement across the highlighted mutation pairs. These examples suggest that small local sequence changes can measurably alter the ranking outputs of the computational framework.</p><p>This study shows that machine learning can prioritize and refine candidate CDR sequences for Creutzfeldt-Jakob disease despite extremely limited training data. Both predictive classifiers distinguished PrP-targeting exemplars from unrelated background sequences using only sequence-derived features. The neutralization-like model achieved strong discriminative performance, while the selectivity-like classifier demonstrated moderate yet statistically significant predictive capability. These results indicate that meaningful sequence-level signal is encoded within antibody sequence features such as residue composition, charge distribution, hydrophobicity, and short motif content.</p><p>Composite ranking across three independent computational screening constraints, including BBB compatibility proxy, anti-PrP selectivity-like scoring, and neutralization potential, allowed the pipeline to approximate candidate-prioritization trade-offs (Dobson et al. 2016; Mieczkowski et al. 2023). Rather than optimizing a single metric, the framework balances multiple biological requirements simultaneously. This multi-objective scoring approach enabled identification of candidates that perform well across all criteria rather than excelling in only one dimension.</p><p>External control comparison using FDA-approved CNS antibodies provided additional evidence that the models capture features associated with PrP-targeting rather than generic CNS-associated properties. These control antibodies scored near zero in both predictive models, while PrP-targeting exemplars clustered near the highest score ranges. This separation supports that the classifiers learned biologically relevant signal associated with PrP targeting rather than simply recognizing general CNS antibody traits.</p><p>The mutation model further expanded the search space by generating targeted sequence modifications around top-ranking candidates. Several mutation-derived variants achieved scores comparable to or exceeding their parent sequences, demonstrating the potential for computational prioritization of candidate antibody sequences. Traditional anti-PrP drug discovery is bottlenecked by the scarcity of validated antibody datasets and the difficulty of experimental prion screening. This pipeline demonstrates that sequence-level machine learning can extract meaningful biological signal from as few as 21 positive examples, but the findings remain preliminary and computational. The identified candidates represent a prioritized set for later validation, including molecular docking against PrP structures, binding affinity assays, cell-based BBB transcytosis models, and prion cell model neutralization testing. Future work would center around gathering more positive exemplars, wet-lab validation, expanding the mutation model, and replacing the BBB proxy with a trained transport model. These results should be interpreted as hypothesis for sequence ranking rather than proof of native PrP<sup>Sc</sup>-specific binding, defined PrP epitope recognition, or BBB permeation. These results should also be interpreted in light of prior reports of anti-PrP antibody-associated neurotoxicity.</p>","references":[{"reference":"Aguzzi A, Calella AM. 2009. Prions: Protein aggregation and infectious diseases. Physiological Reviews. 89: 1105.","pubmedId":"","doi":"10.1152/physrev.00006.2009"},{"reference":"Collinge J, Hawke S. 2009. Prion inhibition.","pubmedId":"","doi":""},{"reference":"Dobson CL, Devine PW, Phillips JJ, Higazi DR, Lloyd C, Popplewell AG. 2016. Engineering the surface properties of a human monoclonal antibody prevents self-association and rapid clearance in vivo. Scientific Reports. 6: 38644.","pubmedId":"","doi":"10.1038/srep38644"},{"reference":"<p>Frontzek K, Bardelli M, Senatore A, Henzi A, Reimann RR, Bedir S, et al., Aguzzi A. 2022. A conformational switch controlling the toxicity of the prion protein. Nat Struct Mol Biol 29(8): 831-840.</p>","pubmedId":"35948768","doi":""},{"reference":"<p>Mercer RCC, Harris DA. 2023. Mechanisms of prion-induced toxicity. Cell Tissue Res 392(1): 81-96.</p>","pubmedId":"36070155","doi":""},{"reference":"Mieczkowski C, Zhang X, Lee D, Nguyen K, Lv W, Wang Y, et al., Gries JM. 2023. Blueprint for antibody biologics developability. mAbs. 15: 2185924.","pubmedId":"","doi":"10.1080/19420862.2023.2185924"},{"reference":"Pankiewicz JE, Sanchez S, Kirshenbaum K, Kascsak RB, Kascsak RJ, Sadowski MJ. 2019. Anti-prion protein antibody 6D11 restores cellular proteostasis of prion protein through disrupting recycling propagation of PrPSc and targeting PrPSc for lysosomal degradation. Molecular Neurobiology. 56: 2073.","pubmedId":"","doi":"10.1007/s12035-018-1208-4"},{"reference":"Prusiner SB. 1998. Prions. Proceedings of the National Academy of Sciences. 95: 13363.","pubmedId":"","doi":"10.1073/pnas.95.23.13363"},{"reference":"<p>Reimann RR, Sonati T, Hornemann S, Herrmann US, Arand M, Hawke S, Aguzzi A. 2016. Differential Toxicity of Antibodies to the Prion Protein. PLoS Pathog 12(1): e1005401.</p>","pubmedId":"26821311","doi":""},{"reference":"<p>Ruiz-López E, Schuhmacher AJ. 2021. Transportation of Single-Domain Antibodies through the Blood-Brain Barrier. Biomolecules 11(8): 10.3390/biom11081131.</p>","pubmedId":"34439797","doi":""},{"reference":"<p>Sonati T, Reimann RR, Falsig J, Baral PK, O'Connor T, Hornemann S, et al., Aguzzi A. 2013. The toxicity of antiprion antibodies is mediated by the flexible tail of the prion protein. Nature 501(7465): 102-6.</p>","pubmedId":"23903654","doi":""},{"reference":"Triguero D, Buciak JB, Pardridge WM. 1989. Blood-brain barrier transport of cationized immunoglobulin G. Proceedings of the National Academy of Sciences. 86: 4761.","pubmedId":"","doi":"10.1073/pnas.86.12.4761"},{"reference":"Uger MD, Chai V, Ciolfi V, Cashman NR, Tian B, Wong WY, Chao HLM. 2016. Antibodies and conjugates that target misfolded prion protein.","pubmedId":"","doi":""},{"reference":"Williamson RA, Burton DR, Prusiner SB. 2001. Antibodies specific for native PrPSc.","pubmedId":"","doi":""},{"reference":"<p>Wu B, McDonald AJ, Markham K, Rich CB, McHugh KP, Tatzelt J, et al., Harris DA. 2017. The N-terminus of the prion protein is a toxic effector regulated by the C-terminus. Elife 6: pii: e23473. 10.7554/eLife.23473.</p>","pubmedId":"28527237","doi":""},{"reference":"Yang W. 2014. Efficient penetration of the blood-brain barrier by anti-prion antibody fragments. Science Translational Medicine. 6: 234ra60.","pubmedId":"","doi":"10.1126/scitranslmed.3008870"}],"title":"<p>Against All Odds: Computational Screening Via Machine Learning Ranking and Generation of Antibody Candidates for Creutzfeldt-Jakob Disease</p>","reviews":[],"curatorReviews":[]},{"id":"e46cabc4-2880-4189-b524-0a23dd12cd01","decision":"publish","abstract":"<p>Creutzfeldt-Jakob disease (CJD) is a fatal prion disorder with no approved treatments. This study developed a machine learning pipeline trained on 21 literature-curated PrP-targeting CDR sequences to rank 29,574 original and mutation-generated candidate sequences using neutralization, selectivity, and literature-based blood-brain barrier (BBB) proxy scores. The models showed internal performance (neutralization AUC = 0.9239; selectivity AUC = 0.7759), identified 10 high-percentile candidates, and generated 7,309 novel variants. These findings support hypothesis-generating sequence-level prioritization of PrP-targeting candidates, but do not establish native PrP<sup>Sc</sup>-specific binding, full-antibody efficacy, exact PrP epitope recognition, or in vivo BBB penetration and require further experimental validation and testing.</p>","acknowledgements":"<p>The authors would like to thank Academy of the Canyons for providing resources to enable this project.</p>","authors":[{"affiliations":["College of the Canyons, Santa Clarita, California, United States","Academy of the Canyons, Santa Clarita, California, United States"],"departments":["",""],"credit":["conceptualization","dataCuration","formalAnalysis","fundingAcquisition","investigation","methodology","project","resources","software","validation","visualization","writing_originalDraft","writing_reviewEditing"],"email":"vtiruvellore@my.canyons.edu","firstName":"Vinayan","lastName":"Tiruvellore","submittingAuthor":true,"correspondingAuthor":true,"equalContribution":false,"WBId":null,"orcid":"https://orcid.org/0009-0003-5451-2503"},{"affiliations":["Academy of the Canyons, Santa Clarita, California, United States","College of the Canyons, Santa Clarita, California, United States"],"departments":["",""],"credit":["supervision","resources"],"email":"mkoegle@hartdistrict.org","firstName":"Mike","lastName":"Koegle","submittingAuthor":false,"correspondingAuthor":false,"equalContribution":false,"WBId":null,"orcid":null}],"awards":[],"conflictsOfInterest":"<p>The authors declare that there are no conflicts of interest present.</p>","dataTable":{"url":null},"extendedData":[],"funding":"<p>This work did not receive external funding. Support was provided by the authors and Academy of the Canyons, Santa Clarita, CA.</p>","image":{"url":"https://portal.micropublication.org/uploads/7a2ba7c6672cc12334c9fef3fb7eed6f.jpg"},"imageCaption":"<p>A) Machine learning pipeline for sequence-level computational prioritization of PrP-targeting CDR candidates. flow chart overview of the computational workflow used to evaluate CDR sequences against three screening criteria: blood-brain barrier compatibility proxy, selectivity-like scoring, and predicted neutralization-like activity. Background antibody sequences were scored relative to PrP-targeting exemplars, ranked by composite candidate-prioritization score, and used as the basis for mutation-based generation of novel candidate variants.</p><p>B) Feature engineering and sequence-level model inputs. Quantitative physicochemical descriptors were extracted from antibody CDR sequences, including length, hydropathy, residue composition, charge distribution, aromatic fraction, and glycine/proline content. These engineered sequence features were used to train classifiers that distinguish PrP-targeting exemplar sequences from non-PrP background sequences and support multi-objective ranking.</p><p>C) Model performance and candidate ranking results. Predicted neutralization-like and selectivity-like scores separated curated PrP-targeting positive exemplars from background antibodies and FDA-approved CNS antibody controls. ROC AUC and Mann-Whitney U test results support statistically significant discrimination between PrP-targeting exemplar sequences and non-PrP comparison groups.</p><p>D) Structural comparison of parent and mutation-derived antibody candidates. ESMFold-predicted CDR structures show conformational effects of single-residue deletions in high-scoring mutation-derived variants. Removal of structurally restrictive residues altered loop geometry and corresponded with improved composite computational ranking relative to the parent sequences.</p>","imageTitle":"<p>Machine Learning-Based Sequence Prioritization and Generation of Candidate Antibody CDR Sequences</p>","methods":"<p>Data collection: This study was designed as a preliminary computational screening pipeline for candidate PrP-targeting CDR sequences rather than as evidence of therapeutic efficacy. The positive set consisted of 21 curated PrP-targeting CDR sequences compiled from patents and literature-derived criteria sets: 8 sequences in the neutralization set, 11 in the selectivity set, and 2 BBB-related comparator sequences. The modeled sequences represent variable-length CDR segments ranging from 5 to 17 amino acids and therefore capture only CDR-derived subregions of the antibody variable domain rather than a complete antibody molecule. Because experimentally validated anti-PrP antibody datasets are scarce, this curated set was used as a small hypothesis-generating positive reference set. For model contrast, more than 20,000 non-PrP human IGH background CDR sequences were used as negatives during training. Additionally, the source reference, sequence identity, exact sequence used, region, and model role for each exemplar are provided in exemplar_table_full.csv in the project GitHub repository.</p><p>Control construction: The primary negative control was the non-PrP IGH background set. An external comparator set consisting of 8 FDA-approved CNS-related antibodies that were not designed to bind PrP was used during downstream evaluation. These FDA CNS antibodies were treated as non-PrP comparators to test whether the scoring framework was capturing PrP-related sequence patterns rather than general CNS-targeting characteristics.</p><p>Feature engineering: Each CDR sequence was converted into a hybrid feature representation consisting of character-level 3-mer sequence features and 7 physicochemical features. The physicochemical features were length, hydrophobic fraction, aromatic fraction, positive-residue fraction, negative-residue fraction, net charge per residue, and glycine/proline fraction. In the final target feature matrix, the 21 positive sequences were represented by 137 distinct 3-mer features and 7 physicochemical features. This representation was chosen to preserve short sequence motif information while also capturing coarse biochemical properties relevant to antibody recognition and developability.</p><p>Machine-learning approach: Two binary classifiers were trained: a neutralization-like classifier and a selectivity-like classifier. Both models used L2-regularized logistic regression with class balancing and the liblinear solver. Positive labels were assigned from the curated PrP-targeting sequence set, and negative labels were assigned from the non-PrP background set. The models therefore estimated whether a sequence more closely resembled the curated neutralization or selectivity exemplars than the background repertoire. The computational framework was trained on sequence-derived features from curated PrP-targeting CDR exemplars and was not trained directly against the full human PrP<sup>C</sup> sequence. It does not predict a specific PrP binding epitope. These outputs should therefore be interpreted as ranking scores rather than direct predictions of full-antibody therapeutic activity or native PrP<sup>Sc</sup>-specific binding.</p><p>Training and validation: Model performance was evaluated using 5-fold stratified cross-validation with ROC AUC and average precision. The neutralization-like classifier achieved mean ROC AUC = 0.9239 ± 0.0810, and the selectivity-like classifier achieved mean ROC AUC = 0.7759 ± 0.1217 across folds. Because only 21 curated positive exemplars were available, these models should be interpreted as preliminary ranking models rather than definitive predictive classifiers. The larger variance observed for the selectivity model is consistent with reduced stability under limited positive-sample conditions.</p><p>BBB proxy scoring model: The BBB proxy scoring model was a hand-engineered sequence-ranking function rather than an experimentally trained BBB penetration model. For each antibody sequence, the model computed six physicochemical descriptors: sequence length, approximate molecular weight, net charge at pH 7.4, estimated isoelectric point (pI), GRAVY hydropathy (Kyte-Doolittle average hydropathy), and a hydrophobic patch score defined as the fraction of 6-residue windows containing at least 4 hydrophobic residues. The proxy then assigned a BBB compatibility score by rewarding sequences with net charge near 6.0 and pI near 9.6 using Gaussian-like soft-band terms, while penalizing excessive charge, excessively high pI, large molecular weight, high hydropathy, and high hydrophobic patchiness using smooth soft-penalty terms. The raw score was computed as:</p><p>0.35 x charge_reward + 0.25 x pI_reward - 0.15 x size_penalty - 0.15 x gravy_penalty - 0.10 x patch_penalty - 0.10 x charge_extreme_penalty - 0.10 x pI_extreme_penalty</p><p>and then transformed to a 0-1 scale with the logistic function where BBB_score = 1 / (1 + e<sup>(-4 * (raw - 0.3)</sup>)). Higher scores indicated stronger predicted BBB compatibility under this proxy, whereas lower scores indicated weaker compatibility. Because no independent BBB transcytosis or transport training dataset was used, this score should be interpreted strictly as a computational ranking proxy for sequence-level BBB compatibility rather than as direct evidence that intact antibodies would penetrate the BBB in vivo.</p><p>Statistical adequacy and overfitting: The positive reference set was small, so overfitting risk is an important limitation. To mitigate this, the classifiers used regularized logistic regression, class balancing, and 5-fold stratified cross-validation, and were evaluated against both a large background repertoire and an external FDA CNS comparator set. These steps support the use of the framework for preliminary prioritization, but they do not eliminate the need for independent experimental validation. Accordingly, all outputs should be interpreted as hypothesis-generating candidate rankings rather than validated therapeutic predictions.</p><p>Candidate ranking: Candidate sequences were ranked by combining the neutralization-like score, the selectivity-like score, and the BBB proxy score. Ranking was performed on original seed sequences and mutation-generated variants. Across 29,574 evaluated sequences in the ranking/validation workflow, the top candidates were prioritized based on combined percentile performance across the three criteria.</p><p>Statistical testing: Mann-Whitney U tests were used to compare score distributions between predefined groups, including positives versus background and positives versus FDA CNS comparators, for both neutralization-like and selectivity-like scores. These tests were used as score-separation analyses rather than proof of causal biological discrimination or native PrP<sup>Sc</sup>-specific binding.</p><p>Mutation modeling: Sequence diversification was performed computationally by targeted mutation and rescoring. This procedure generated 7,309 novel sequences not present in the original library. These mutated candidates were evaluated using the same ranking framework as the original sequences.</p><p>Data and code availability: The curated input datasets, extracted feature tables, trained scoring artifacts, ranked candidate outputs, analysis code, and the exemplar table with the CDR sequences the model was trained on are publicly available at https://github.com/Axy-lis/CJD_ML_AAO.</p>","reagents":"<p></p>","patternDescription":"<p>CJD is a rare but universally fatal prion disease. There is no cure, no disease-modifying therapy. CJD is caused by the misfolding of native prion protein, PrP<sup>C</sup>, into the pathogenic isoform PrP<sup>Sc</sup>. PrP<sup>Sc</sup> acts as a self-propagating template, converts normal proteins into misfolded forms, forms amyloid aggregates, and causes widespread neuronal death (Prusiner 1998; Aguzzi and Calella 2009). Therapeutic development is challenging because the pathogenic PrP<sup>Sc</sup> fold can differ substantially between species and disease-associated prion strains, so compounds effective against mouse prions may fail against human prions. Additionally, the blood-brain barrier (BBB) limits CNS delivery, and therapies must avoid disrupting normal PrP<sup>C</sup> function. Despite decades of research, no clinically approved antibody or small-molecule therapy exists (Aguzzi and Calella 2009; Pankiewicz et al. 2019).</p><p>Many reported anti-PrP antibodies act through binding to PrP<sup>C</sup> or related conformational states, and apparent PrP<sup>Sc</sup> recognition can depend on denaturation or assay-specific exposure of epitopes. Accordingly, the present study should not be interpreted as establishing selective binding to native PrP<sup>Sc</sup>. In addition, prior studies have reported neurotoxicity associated with some anti-PrP antibodies and related PrP toxic signaling pathways, emphasizing that antibody-based intervention in prion disease remains biologically complex and that the present computational rankings should not be interpreted as evidence of safety or therapeutic suitability (Sonati et al. 2013; Reimann et al. 2016; Wu et al. 2017; Frontzek et al. 2022; Mercer and Harris 2023).</p><p>Machine learning approaches to antibody engineering are promising in identifying candidate antibodies. By training ranking models on sequence-derived CDR features from curated anti-PrP antibody exemplars, it is possible to evaluate large candidate libraries and identify those most likely to satisfy computational screening criteria. The experimental question was whether a machine learning-driven computational pipeline could prioritize candidate PrP-targeting CDR segments for CJD using neutralization-like, selectivity-like, and BBB compatibility proxy criteria. The hypothesis was that a machine learning pipeline trained on sequence-derived features from literature-curated PrP-targeting CDR exemplars would rank anti-PrP-like candidates above background sequences, distinguish them from unrelated CNS antibodies, and generate novel, high-scoring variants through mutation modeling. The modeled inputs in this study were variable-length CDR segments rather than complete antibody molecules or fixed-length 10-mers. In the curated exemplar set, these CDR sequences ranged from 5 to 17 amino acids, with many clustering near approximately 10 amino acids. The framework was therefore designed for sequence-level prioritization rather than full-antibody functional prediction.</p><p>The primary control was baseline background antibodies. An external negative control used FDA-approved CNS-targeting antibodies to test whether the model captured anti-PrP sequence signatures rather than general CNS-targeting characteristics. For data collection, 21 literature-curated PrP-targeting CDR exemplars were used as positives, and more than 20,000 non-PrP sequences were used as background. Anti-PrP antibody examples and PrP-targeting antibody concepts were drawn from prior literature and patent sources (Pankiewicz et al. 2019; Uger et al. 2016; Williamson et al. 2001; Collinge and Hawke 2009). Sequences derived from Uger et al. 2016 were included as literature-reported PrP-targeting exemplars for computational comparison, but not as definitive evidence of selective native PrP<sup>Sc</sup> recognition. For feature engineering, CDR sequences were converted into quantitative descriptors, including 3-mer motif embeddings, amino acid composition, length, net charge at pH 7.4, hydropathy using the Kyte-Doolittle scale, aromatic fraction, and residue class proportions.</p><p>Model training was designed to distinguish positives from background. The pipeline included a neutralization-like classifier, a selectivity-like classifier, and a BBB compatibility proxy. Candidate ranking was performed by scoring more than 20,000 sequences and calculating a composite score across all three criteria. The top 25 candidates were selected. Mutation modeling was then applied through targeted sequence modifications, and modified sequences were re-scored to identify novel high-ranking candidates.</p><p>The neutralization-like classifier achieved ROC AUC = 0.9239 ± 0.0810, and the selectivity-like classifier achieved ROC AUC = 0.7759 ± 0.1217. Both were above random expectation, supporting that CDR sequence-derived features carry meaningful signal for PrP-targeting exemplar ranking. The BBB compatibility proxy was a literature-based heuristic informed by papers on antibody BBB transport and anti-PrP antibody fragments (Triguero et al. 1989; Yang et al. 2014; Ruiz-López et al. 2021). Strong internal model performance was achieved despite training on only 21 curated exemplars.</p><p>Mann-Whitney U tests showed statistically significant score separation between positive exemplars and background sequences and between positives and FDA CNS antibodies. Positives versus background for neutralization had a U statistic of 180000 and a p-value of 0.000000102. Positives versus background for selectivity had a U statistic of 300000 and a p-value of 0.0000000001. Positives versus FDA CNS antibodies for neutralization had a U statistic of 72 and a p-value of 0.0003129. Positives versus FDA CNS antibodies for selectivity had a U statistic of 120 and a p-value of 0.0000609. These low p-values support score separation between predefined groups, but do not independently establish antibody activity, selective native PrP<sup>Sc</sup> recognition, or guarantee model generalizability.</p><p>The neutralization score distributions showed positive exemplars clustered near the highest predicted values, around 0.99 to 1.0, whereas background sequences remained spread across lower score values, around 0.3 to 0.7. The selectivity-like score distribution showed a similar pattern, with positive exemplars again occupying the highest scoring region and background sequences occupying lower score ranges. Neutralization-like and selectivity-like scores were also compared across known PrP-targeting exemplars, FDA CNS antibodies, and background sequences. Positive exemplars clustered near maximal scores, FDA CNS antibodies scored near zero, and background sequences showed low-score distributions. Together, these comparisons support that the classifiers captured features associated with the PrP-targeting reference set rather than general antibody features.</p><p>The ranked candidate pool consisted of more than 20,000 original antibody sequences and about 7,000 mutation-generated variants. This search space allowed the machine learning models to identify rare high-scoring candidates within diverse antibody sequence space. Ten candidates ranked in the 95th to 99th percentile across 29,574 evaluated sequences. Positive exemplars were heavily enriched in the top 50, while FDA CNS antibodies were entirely absent. Mutation modeling generated 7,309 novel sequences not present in the original library. The top novel candidates ranked 3rd, 7th, and 12th overall, outperforming 99.97% of all evaluated sequences in this computational screen.</p><p>Single-residue deletions produced the largest improvements. In one case, parent sequence QQNKNWPPGT ranked 17, while novel sequence QQNKNWPGT ranked 3. A single Pro8 deletion distinguished the parent from the novel generated sequence. Removing Pro8 relieved local conformational constraint in the CDR-segment loop and improved the composite computational score from 0.9738 to 0.9810, advancing the sequence from rank 17 to rank 3 out of 29,574 total candidates. In another case, parent sequence SSYTITNTQK ranked 367, while novel sequence SSTITNTQK ranked 12. A single Tyr3 deletion distinguished the parent from the novel generated sequence and increased the combined score from 0.9439 to 0.9759, producing the largest score improvement across the highlighted mutation pairs. These examples suggest that small local sequence changes can measurably alter the ranking outputs of the computational framework.</p><p>This study shows that machine learning can prioritize and refine candidate CDR sequences for Creutzfeldt-Jakob disease despite extremely limited training data. Both predictive classifiers distinguished PrP-targeting exemplars from unrelated background sequences using only sequence-derived features. The neutralization-like model achieved strong discriminative performance, while the selectivity-like classifier demonstrated moderate yet statistically significant predictive capability. These results indicate that meaningful sequence-level signal is encoded within antibody sequence features such as residue composition, charge distribution, hydrophobicity, and short motif content.</p><p>Composite ranking across three independent computational screening constraints, including BBB compatibility proxy, anti-PrP selectivity-like scoring, and neutralization potential, allowed the pipeline to approximate candidate-prioritization trade-offs (Dobson et al. 2016; Mieczkowski et al. 2023). Rather than optimizing a single metric, the framework balances multiple biological requirements simultaneously. This multi-objective scoring approach enabled identification of candidates that perform well across all criteria rather than excelling in only one dimension.</p><p>External control comparison using FDA-approved CNS antibodies provided additional evidence that the models capture features associated with PrP-targeting rather than generic CNS-associated properties. These control antibodies scored near zero in both predictive models, while PrP-targeting exemplars clustered near the highest score ranges. This separation supports that the classifiers learned biologically relevant signal associated with PrP targeting rather than simply recognizing general CNS antibody traits.</p><p>The mutation model further expanded the search space by generating targeted sequence modifications around top-ranking candidates. Several mutation-derived variants achieved scores comparable to or exceeding their parent sequences, demonstrating the potential for computational prioritization of candidate antibody sequences. Traditional anti-PrP drug discovery is bottlenecked by the scarcity of validated antibody datasets and the difficulty of experimental prion screening. This pipeline demonstrates that sequence-level machine learning can extract meaningful biological signal from as few as 21 positive examples, but the findings remain preliminary and computational. The identified candidates represent a prioritized set for later validation, including molecular docking against PrP structures, binding affinity assays, cell-based BBB transcytosis models, and prion cell model neutralization testing. Future work would center around gathering more positive exemplars, wet-lab validation, expanding the mutation model, and replacing the BBB proxy with a trained transport model. These results should be interpreted as hypothesis for sequence ranking rather than proof of native PrP<sup>Sc</sup>-specific binding, defined PrP epitope recognition, or BBB permeation. These results should also be interpreted in light of prior reports of anti-PrP antibody-associated neurotoxicity.</p>","references":[{"reference":"Aguzzi A, Calella AM. 2009. Prions: Protein aggregation and infectious diseases. Physiological Reviews. 89: 1105.","pubmedId":"","doi":"10.1152/physrev.00006.2009"},{"reference":"Collinge J, Hawke S. 2009. Prion inhibition.","pubmedId":"","doi":""},{"reference":"Dobson CL, Devine PW, Phillips JJ, Higazi DR, Lloyd C, Popplewell AG. 2016. Engineering the surface properties of a human monoclonal antibody prevents self-association and rapid clearance in vivo. Scientific Reports. 6: 38644.","pubmedId":"","doi":"10.1038/srep38644"},{"reference":"<p>Frontzek K, Bardelli M, Senatore A, Henzi A, Reimann RR, Bedir S, et al., Aguzzi A. 2022. A conformational switch controlling the toxicity of the prion protein. Nat Struct Mol Biol 29(8): 831-840.</p>","pubmedId":"35948768","doi":""},{"reference":"<p>Mercer RCC, Harris DA. 2023. Mechanisms of prion-induced toxicity. Cell Tissue Res 392(1): 81-96.</p>","pubmedId":"36070155","doi":""},{"reference":"Mieczkowski C, Zhang X, Lee D, Nguyen K, Lv W, Wang Y, et al., Gries JM. 2023. Blueprint for antibody biologics developability. mAbs. 15: 2185924.","pubmedId":"","doi":"10.1080/19420862.2023.2185924"},{"reference":"Pankiewicz JE, Sanchez S, Kirshenbaum K, Kascsak RB, Kascsak RJ, Sadowski MJ. 2019. Anti-prion protein antibody 6D11 restores cellular proteostasis of prion protein through disrupting recycling propagation of PrPSc and targeting PrPSc for lysosomal degradation. Molecular Neurobiology. 56: 2073.","pubmedId":"","doi":"10.1007/s12035-018-1208-4"},{"reference":"Prusiner SB. 1998. Prions. Proceedings of the National Academy of Sciences. 95: 13363.","pubmedId":"","doi":"10.1073/pnas.95.23.13363"},{"reference":"<p>Reimann RR, Sonati T, Hornemann S, Herrmann US, Arand M, Hawke S, Aguzzi A. 2016. Differential Toxicity of Antibodies to the Prion Protein. PLoS Pathog 12(1): e1005401.</p>","pubmedId":"26821311","doi":""},{"reference":"<p>Ruiz-López E, Schuhmacher AJ. 2021. Transportation of Single-Domain Antibodies through the Blood-Brain Barrier. Biomolecules 11(8): 10.3390/biom11081131.</p>","pubmedId":"34439797","doi":""},{"reference":"<p>Sonati T, Reimann RR, Falsig J, Baral PK, O'Connor T, Hornemann S, et al., Aguzzi A. 2013. The toxicity of antiprion antibodies is mediated by the flexible tail of the prion protein. Nature 501(7465): 102-6.</p>","pubmedId":"23903654","doi":""},{"reference":"Triguero D, Buciak JB, Pardridge WM. 1989. Blood-brain barrier transport of cationized immunoglobulin G. Proceedings of the National Academy of Sciences. 86: 4761.","pubmedId":"","doi":"10.1073/pnas.86.12.4761"},{"reference":"Uger MD, Chai V, Ciolfi V, Cashman NR, Tian B, Wong WY, Chao HLM. 2016. Antibodies and conjugates that target misfolded prion protein.","pubmedId":"","doi":""},{"reference":"Williamson RA, Burton DR, Prusiner SB. 2001. Antibodies specific for native PrPSc.","pubmedId":"","doi":""},{"reference":"<p>Wu B, McDonald AJ, Markham K, Rich CB, McHugh KP, Tatzelt J, et al., Harris DA. 2017. The N-terminus of the prion protein is a toxic effector regulated by the C-terminus. Elife 6: pii: e23473. 10.7554/eLife.23473.</p>","pubmedId":"28527237","doi":""},{"reference":"Yang W. 2014. Efficient penetration of the blood-brain barrier by anti-prion antibody fragments. Science Translational Medicine. 6: 234ra60.","pubmedId":"","doi":"10.1126/scitranslmed.3008870"}],"title":"<p>Against All Odds: Computational Screening Via Machine Learning Ranking and Generation of Antibody Candidates for Creutzfeldt-Jakob Disease</p>","reviews":[],"curatorReviews":[]}]}},"species":{"species":[{"value":"acer saccharum","label":"Acer saccharum","imageSrc":"","imageAlt":"","mod":"TreeGenes","modLink":"https://treegenesdb.org","linkVariable":""},{"value":"achillea millefolium","label":"Achillea millefolium","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"acinetobacter baylyi","label":"Acinetobacter baylyi","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"actinobacteria bacterium","label":"Actinobacteria bacterium","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"adelges tsugae","label":"Adelges tsugae","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"adenocaulon chilense","label":"Adenocaulon chilense","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"aedes japonicus","label":"Aedes japonicus","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"aegorhinus vitulus","label":"Aegorhinus vitulus","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"alaimidae","label":"Alaimidae","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"allobates femoralis","label":"Allobates femoralis","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"alnus glutinosa","label":"Alnus glutinosa","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"alosa aestivalis","label":"Alosa aestivalis","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"alosa pseudoharengus","label":"Alosa pseudoharengus","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"alternaria alternata","label":"Alternaria alternata","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"amynthas agrestis","label":"Amynthas Agrestis","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"ancylostoma caninum","label":"Ancylostoma caninum","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"ancylostoma ceylanicum","label":"Ancylostoma ceylanicum","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"anemone multifida","label":"Anemone multifida","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"anguilla rostrata","label":"Anguilla rostrata","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"anisakis simplex","label":"Anisakis simplex","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"anomala albopilosa","label":"Anomala albopilosa","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"anthomyiidae sp","label":"Anthomyiidae sp","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"anthomyiidae sp","label":"Anthomyiidae sp","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"arabidopsis","label":"Arabidopsis","imageSrc":"arabidopsis.png","imageAlt":"Arabidopsis graphic by Zoe Zorn CC BY 4.0","mod":"TAIR","modLink":"https://arabidopsis.org","linkVariable":""},{"value":"architeuthis dux","label":"Architeuthis dux","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"arion vulgaris","label":"Arion vulgaris","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"armeria","label":"Armeria","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"artemia","label":"Artemia","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"arthrobacter sp.","label":"Arthrobacter sp.","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"ascaridia","label":"Ascaridia","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"ascaridia galli","label":"Ascaridia galli","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"asparagopsis taxiformis","label":"Asparagopsis taxiformis","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"astatotilapia burtoni","label":"Astatotilapia burtoni","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"avena sativa","label":"Avena sativa","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"aves","label":"Aves","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"bacillus","label":"Bacillus (firmicutes)","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"bacillus cereus","label":"Bacillus cereus","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"bacillus mycoides","label":"Bacillus mycoides","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"bacillus subtilis","label":"Bacillus subtilis","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"bacillus thuringiensis","label":"Bacillus thuringiensis","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"bacillus toyonensis","label":"Bacillus toyonensis","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"bacillus wiedmannii","label":"Bacillus wiedmannii","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"bacteria","label":"Bacteria","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"bacteriophage","label":"Bacteriophage","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"bactrocera","label":"Bactrocera sp.","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"batrachospermum gelatinosum","label":"Batrachospermum gelatinosum","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"betula lenta","label":"Betula lenta","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"betula nigra","label":"Betula nigra","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"bombus dahlbohmii","label":"Bombus dahlbohmii","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"bombus terrestris","label":"Bombus terrestris","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"bombyx mori","label":"Bombyx mori","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"bos taurus","label":"Bos Taurus","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"brachygobius doriae","label":"Brachygobius doriae","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"brassica oleracea","label":"Brassica oleracea","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"brassica rapa","label":"Brassica rapa","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"brugia malayi","label":"Brugia malayi","imageSrc":"","imageAlt":"","mod":"WormBase","modLink":"www.wormbase.org","linkVariable":""},{"value":"burkholderia thailandensis","label":"Burkholderia thailandensis","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"buttiauxella","label":"Buttiauxella","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"caenorhabditis brenneri","label":"Caenorhabditis brenneri","imageSrc":"","imageAlt":"","mod":"WormBase","modLink":"www.wormbase.org","linkVariable":""},{"value":"caenorhabditis briggsae","label":"Caenorhabditis briggsae","imageSrc":"","imageAlt":"","mod":"WormBase","modLink":"www.wormbase.org","linkVariable":""},{"value":"c. elegans","label":"Caenorhabditis elegans","imageSrc":"c-elegans.jpg","imageAlt":"C. elegans graphic by Zoe Zorn CC BY 4.0","mod":"WormBase","modLink":"https://wormbase.org","linkVariable":""},{"value":"caenorhabditis inopinata","label":"Caenorhabditis inopinata","imageSrc":"","imageAlt":"","mod":"WormBase","modLink":"www.wormbase.org","linkVariable":""},{"value":"caenorhabditis japonica","label":"Caenorhabditis japonica","imageSrc":"","imageAlt":"","mod":"WormBase","modLink":"www.wormbase.org","linkVariable":""},{"value":"caenorhabditis nigoni","label":"Caenorhabditis nigoni","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"caenorhabditis remanei","label":"Caenorhabditis remanei","imageSrc":"","imageAlt":"","mod":"WormBase","modLink":"www.wormbase.org","linkVariable":""},{"value":"caenorhabditis tropicalis","label":"Caenorhabditis tropicalis","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"calidifontibacillus","label":"Calidifontibacillus","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"calidifontibacillus erzuremensis","label":"Calidifontibacillus erzuremensis","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"calliphora sp","label":"Calliphora sp","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"caltha sagittata","label":"Caltha sagittata","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"cambarus latimanus","label":"Cambarus latimanus","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"candida albicans","label":"Candida albicans","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"canis familiaris","label":"Canis familiaris","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"cannabis sativa","label":"Cannabis sativa","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"caretta caretta","label":"Caretta caretta","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"cassiopea xamachana","label":"Cassiopea xamachana","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"caulobacter vibrioides","label":"Caulobacter vibrioides","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"cephalopods","label":"Cephalopoda","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"cerastium arvense","label":"Cerastium arvense","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"ceriodaphnia","label":"Ceriodaphnia","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"ceroglossus suturalis","label":"Ceroglossus suturalis","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"chaetoceros","label":"Chaetoceros","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"chamaecrista fasciculata","label":"Chamaecrista fasciculata","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"chilicola chalcidiformis","label":"Chilicola chalcidiformis","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"chitinimonas","label":"Chitinimonas","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"chlamydomonas reinhardtii","label":"Chlamydomonas reinhardtii","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"chromobacterium","label":"Chromobacterium","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"chrysemys picta","label":"Chrysemys picta","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"chrysoperla rufilabris","label":"Chrysoperla rufilabris","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"citrus","label":"Citrus","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"clavibacter sp.","label":"Clavibacter sp.","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"colinus virginianus","label":"Colinus virginianus","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"crassostrea virginica","label":"Crassostrea virginica","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"crithidia fasciculata","label":"Crithidia fasciculata","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"cutibacterium acnes","label":"Cutibacterium acnes","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"cyanobacteria","label":"Cyanobacteria","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"daphnia","label":"Daphnia","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"daphnia pulex","label":"Daphnia pulex","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"dermacoccus nishinomiyaensis","label":"Dermacoccus nishinomiyaensis","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"diabrotica virgifera","label":"Diabrotica virgifera","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"diabrotica virgifera virgifera virus 1","label":"Diabrotica virgifera virgifera virus 1","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"d. discoideum","label":"Dictyostelium discoideum","imageSrc":"dicty.png","imageAlt":"D. discoideum","mod":"dictyBase","modLink":"http://dictybase.org","linkVariable":""},{"value":"diptera","label":"Diptera","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"dotocryptus bellicosus","label":"Dotocryptus bellicosus","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"drechmeria coniospora","label":"Drechmeria coniospora","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"drosophila","label":"Drosophila","imageSrc":"drosophila.png","imageAlt":"Drosophila graphic by Zoe Zorn CC BY 4.0","mod":"FlyBase","modLink":"https://flybase.org/doi/","linkVariable":"doi"},{"value":"dryopteris campyloptera","label":"Dryopteris campyloptera","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"dryopteris expansa","label":"Dryopteris expansa","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"dryopteris intermedia","label":"Dryopteris intermedia","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"dugesia dorotocephala","label":"Dugesia dorotocephala","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"elasmobranchii","label":"Elasmobranchii","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"embryophyta","label":"Embryophyta","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"enoploteuthis chunii","label":"Enoploteuthis chunii","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"enterobacter aerogenes","label":"Enterobacter aerogenes","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"enterococcus raffinosus","label":"Enterococcus raffinosus","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"epichloë coenophiala","label":"Epichloë coenophiala","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"equus caballus","label":"Equus caballus","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"erigeron sp","label":"Erigeron sp","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"eristalis","label":"Eristalis","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"eruca vesicaria","label":"Eruca vesicaria","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"erwinia carotovora","label":"Erwinia carotovora","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"erythronium americanum","label":"Erythronium americanum","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"escherichia coli","label":"Escherichia coli","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"eukaryota","label":"Eukaryotes","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"felis catus","label":"Felis catus","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"francisella novicida","label":"Francisella novicida","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"francisella tularensis","label":"Francisella tularensis","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"fraxinus americana","label":"Fraxinus americana","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"fucus distichus","label":"Fucus distichus","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"fungi","label":"Fungi","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"gasteropelecus sp.","label":"Gasteropelecus sp.","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"geranium sp","label":"Geranium sp","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"girardia","label":"Girardia","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"glaucomys volans","label":"Glaucomys volans","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"glycine max","label":"Glycine max","imageSrc":"","imageAlt":"","mod":"Soybase","modLink":"https://soybase.org","linkVariable":""},{"value":"glyptemys insculpta","label":"Glyptemys insculpta","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"gossypium hirsutum","label":"Gossypium hirsutum","imageSrc":"","imageAlt":"","mod":"CottonGen","modLink":"https://www.cottongen.org/","linkVariable":""},{"value":"gromphadorhina portentosa","label":"Gromphadorhina portentosa","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"gryllodes sigillatus","label":"Gryllodes sigillatus","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"haliotis rufescens","label":"Haliotis rufescens","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"hepacivirus hominis","label":"Hepatitis C Virus","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"herpes simplex virus type 1","label":"Herpes simplex virus type 1","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"human","label":"Human","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"human coronavirus oc43","label":"Human coronavirus OC43","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"hydra vulgaris","label":"Hydra vulgaris","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"hydropsyche sp","label":"Hydropsyche sp","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"hymenoptera","label":"Hymenoptera","imageSrc":"","imageAlt":"","mod":"Hymenoptera Genome Database","modLink":"https://hymenoptera.elsiklab.missouri.edu/","linkVariable":""},{"value":"hypochaeris radicata","label":"Hypochaeris radicata","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"hypodynerus vespiformis","label":"Hypodynerus vespiformis","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"iflaviridae","label":"Iflaviridae","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"iflavuris","label":"Iflavirus","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"ipomoea hederacea","label":"Ipomoea hederacea","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"ischnomera","label":"Ischnomera","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"ischnomera ruficollis","label":"Ischnomera ruficollis","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"julidochromis marlieri","label":"Julidochromis marlieri","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"juniperus virginiana","label":"Juniperus virginiana","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"kluyveromyces marxianus","label":"Kluyveromyces marxianus","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"l. casei","label":"L. casei","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"lacticaseibacillus casei","label":"Lacticaseibacillus casei","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"larentiinae sp","label":"Larentiinae sp","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"laurus nobilis","label":"Laurus nobilis","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"lepidoptera","label":"Lepidoptera","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"leucanthemum vulgare","label":"Leucanthemum vulgare","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"linepithema humile","label":"Linepithema humile","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"liometopum occidentale","label":"Liometopum occidentale","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"lolium arundinaceum","label":"Lolium arundinaceum","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"lontra longicaudis","label":"Lontra longicaudis","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"lumbriculus variegatus","label":"Lumbriculus variegatus","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"lumbricus terrestris","label":"Lumbricus terrestris","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"lupinus polyphyllus","label":"Lupinus polyphyllus","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"lycorma delicatula","label":"Lycorma delicatula","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"lynx rufus","label":"Lynx rufus","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"magnaporthe oryzae","label":"Magnaporthe oryzae","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"mammalia","label":"Mammalia","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"manihot esculenta","label":"Manihot esculenta","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"medicago lupulina","label":"Medicago lupulina","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"meloidogyne","label":"Meloidogyne","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"mimus polyglottos","label":"Mimus polyglottos","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"bryophyta","label":"Mosses","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"mouse","label":"Mouse","imageSrc":"","imageAlt":"","mod":"MGI","modLink":"https://informatics.jax.org","linkVariable":""},{"value":"m. minutoides","label":"Mus minutoides","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"mycobacterium smegmatis","label":"Mycobacterium smegmatis","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"nakaseomyces glabratus","label":"Nakaseomyces glabratus","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"nauphoeta cinerea","label":"Nauphoeta cinerea","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"neurospora","label":"Neurospora","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"n. benthamiana","label":"Nicotiana benthamiana","imageSrc":"","imageAlt":"","mod":"Solgenomics Network","modLink":"https://solgenomics.net/organism/Nicotiana_benthamiana/genome","linkVariable":""},{"value":"nicotiana tabacum","label":"Nicotiana tabacum","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"noctuidae","label":"Noctuidae","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"noctuidae sp","label":"Noctuidae sp","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"nothobranchius furzeri","label":"Nothobranchius furzeri","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"onchocerca volvulus","label":"Onchocerca volvulus","imageSrc":"","imageAlt":"","mod":"WormBase","modLink":"www.wormbase.org","linkVariable":""},{"value":"orconectes virilis","label":"Orconectes virilis","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"ormia ochracea","label":"Ormia ochracea","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"o. sativa","label":"Oryza sativa","imageSrc":"","imageAlt":"","mod":"Gramene","modLink":"https://www.gramene.org/","linkVariable":""},{"value":"other","label":"Other","imageSrc":"","imageAlt":"","mod":null,"modLink":null,"linkVariable":null},{"value":"oxalis enneaphylla","label":"Oxalis enneaphylla","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"paenarthrobacter nicotinovorans","label":"Paenarthrobacter nicotinovorans","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"paenarthrobacter nicotinovorans","label":"Paenarthrobacter nicotinovorans","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"pantoea","label":"Pantoea","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"pantoea agglomerans","label":"Pantoea agglomerans","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"papaver sp","label":"Papaver sp","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"paramecium bursaria","label":"Paramecium bursaria","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"partitiviridae","label":"Partitiviridae","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"pelodiscus sinensis","label":"Pelodiscus sinensis","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"perezia recurvata","label":"Perezia recurvata","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"petromyzon marinus","label":"Petromyzon marinus","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"photinus pyralis","label":"Photinus pyralis","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"photinus pyralis associated partiti-like virus","label":"Photinus pyralis associated partiti-like virus","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"photinus pyralis iflavirus 1","label":"Photinus pyralis iflavirus 1","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"physcomitrium patens","label":"Physcomitrium patens","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"pinus strobus","label":"Pinus strobus","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"pinus taeda","label":"Pinus taeda","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"platycheirus","label":"Platycheirus","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"plectus sambesii","label":"Plectus sambesii","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"pogonomyrmex occidentalis","label":"Pogonomyrmex occidentalis","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"poncirus trifoliata","label":"Poncirus trifoliata","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"populus deltoides","label":"Populus deltoides","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"potato virus y","label":"Potato virus Y","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"primula magellanica","label":"Primula magellanica","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"pristionchus pacificus","label":"Pristionchus pacificus","imageSrc":"","imageAlt":"","mod":"WormBase","modLink":"www.wormbase.org","linkVariable":""},{"value":"prunus persica","label":"Prunus persica","imageSrc":"","imageAlt":"","mod":"Genome Database for Rosaceae","modLink":"https://www.rosaceae.org/","linkVariable":""},{"value":"psalmopoeus iriminia","label":"Psalmopoeus iriminia","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"pseudanabaena sp.","label":"Pseudanabaena sp.","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"pseudomonas","label":"Pseudomonas","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"pseudomonas aeruginosa","label":"Pseudomonas aeruginosa","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"pseudomonas glycinae","label":"Pseudomonas glycinae","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"pseudomonas putida","label":"Pseudomonas putida","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"pseudomonas syringae","label":"Pseudomonas syringae","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"pterophyllum scalare","label":"Pterophyllum scalare","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"python regius","label":"Python regius","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"quercus macrocarpa","label":"Quercus macrocarpa","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"ralstonia solanacearum","label":"Ralstonia solanacearum","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"ranitomeya imitator","label":"Ranitomeya imitator","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"ranunculus peduncularis","label":"Ranunculus peduncularis","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"rat","label":"Rat","imageSrc":"","imageAlt":"","mod":"RGD","modLink":"https://rgd.mcw.edu","linkVariable":""},{"value":"rheinheimera","label":"Rheinheimera","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"ribes rubrum","label":"Ribes rubrum","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"sars-cov-2","label":"SARS-CoV-2","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"s. cerevisiae","label":"Saccharomyces cerevisiae","imageSrc":"yeast.png","imageAlt":"Yeast graphic by Zoe Zorn CC BY 4.0","mod":"SGD","modLink":"https://yeastgenome.org","linkVariable":""},{"value":"saccharomyces paradoxus","label":"Saccharomyces paradoxus ","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"s. uvarum","label":"Saccharomyces uvarum","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"schistosoma","label":"Schistosoma","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"schizosaccharomyces japonicus","label":"Schizosaccharomyces japonicus","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"s. pombe","label":"Schizosaccharomyces pombe","imageSrc":"pombe.png","imageAlt":"Pombe graphic by Zoe Zorn © Caltech","mod":"PomBase","modLink":"https://www.pombase.org/reference/PMID:","linkVariable":"pmId"},{"value":"schmidtea mediterranea","label":"Schmidtea mediterranea","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"senecio sp","label":"Senecio sp","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"simocephalus","label":"Simocephalus","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"siraitia grosvenorii","label":"Siraitia grosvenorii","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"solanum lycopersicum","label":"Solanum lycopersicum","imageSrc":"","imageAlt":"","mod":"Solgenomics Network","modLink":"https://solgenomics.net/organism/1/view/","linkVariable":""},{"value":"sorghum","label":"Sorghum","imageSrc":"","imageAlt":"","mod":"SorghumBase","modLink":"https://www.sorghumbase.org","linkVariable":""},{"value":"spiroplasma eriocheiris","label":"Spiroplasma eriocheiris","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"staphylococcus aureus","label":"Staphylococcus aureus","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"staphylococcus epidermidis","label":"Staphylococcus epidermidis","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"steinernema carpocapsae","label":"Steinernema carpocapsae","imageSrc":"","imageAlt":"","mod":"WormBase","modLink":"https://wormbase.org","linkVariable":""},{"value":"steinernema hermaphroditum","label":"Steinernema hermaphroditum","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"stenotrophomonas geniculata","label":"Stenotrophomonas geniculata","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"stewartia floidana","label":"Stewartia floridana","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"streptococcus gordonii ","label":"Streptococcus gordonii ","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"streptococcus mutans","label":"Streptococcus mutans","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":" streptococcus pneumoniae","label":"Streptococcus pneumoniae","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"s. purpuratus","label":"Strongylocentrotus purpuratus","imageSrc":"","imageAlt":"","mod":"Echinobase","modLink":"https://www.echinobase.org","linkVariable":""},{"value":"strongyloides ratti","label":"Strongyloides ratti","imageSrc":"","imageAlt":"","mod":"WormBase","modLink":"www.wormbase.org","linkVariable":""},{"value":"sulfolobus","label":"Sulfolobus","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"symphoricarpos albus","label":"Symphoricarpos albus","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"syncirsodes","label":"Syncirsodes","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"synechococcus elongatus","label":"Synechococcus elongatus","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"syrphidae","label":"Syrphidae","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"tarantobelus jeffdanielsi","label":"Tarantobelus jeffdanielsi","imageSrc":"","imageAlt":"","mod":"WormBase","modLink":"www.wormbase.org","linkVariable":""},{"value":"taraxacum officinale","label":"Taraxacum officinale","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"tatochila theodice","label":"Tatochila theodice","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"tetrahymena","label":"Tetrahymena","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"tetramorium immigrans","label":"Tetramorium immigrans","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"tomato brown rugose fruit virus","label":"ToBRFV","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"trachemys scripta","label":"Trachemys scripta","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"tribolium castaneum","label":"Tribolium castaneum","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"trichoptera","label":"Trichoptera","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"trichuris muris","label":"Trichuris muris","imageSrc":"","imageAlt":"","mod":"WormBase","modLink":"www.wormbase.org","linkVariable":""},{"value":"trifolium repens","label":"Trifolium repens","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"trypoxylus dichotomus","label":"Trypoxylus dichotomus","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"tsuga canadensis","label":"Tsuga canadensis","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"ulva expansa","label":"Ulva expansa","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"universal","label":"Universal","imageSrc":"","imageAlt":"","mod":null,"modLink":null,"linkVariable":null},{"value":"vargula hilgendorfii","label":"Vargula hilgendorfii","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"vespula vulgaris","label":"Vespula vulgaris","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"virus","label":"Virus","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"watasenia scintillans","label":"Watasenia scintillans","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"wolbachia pipientis","label":"Wolbachia pipientis","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"xenopus","label":"Xenopus","imageSrc":"xenopus.png","imageAlt":"Xenopus graphic by Zoe Zorn CC BY 4.0","mod":"XenBase","modLink":"https://xenbase.org","linkVariable":""},{"value":"xenorhabdus griffiniae","label":"Xenorhabdus griffiniae","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"yramea cytheris","label":"Yramea cytheris","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"zaprionus indianus","label":"Zaprionus indianus","imageSrc":"","imageAlt":"","mod":"","modLink":"","linkVariable":""},{"value":"zea mays","label":"Zea mays","imageSrc":"","imageAlt":"","mod":"MaizeGDB","modLink":"https://www.maizegdb.org","linkVariable":""},{"value":"zebrafish","label":"Zebrafish","imageSrc":"zebrafish.png","imageAlt":"Zebrafish graphic by Zoe Zorn CC BY 4.0","mod":"ZFIN","modLink":"https://zfin.org","linkVariable":""}]}},"pageContext":{"id":"72d47e77-f90e-4828-9273-b88b9ce9969f","citedBy":[],"parsedCsv":{"csvHeader":[],"csvData":[]}}},
    "staticQueryHashes": ["2114697108"]}