euxhen.hasanaj

Publications

2026

EMBO Journal

SenSet defines cell-type specific senescence signatures in the aged human lung

Euxhen Hasanaj, Delphine Beaulieu, Cankun Wang, et al.

abstract

Cellular senescence is defined as an irreversible growth arrest observed when cells are exposed to a variety of stressors, including DNA damage, oxidative stress, or nutrient deprivation. Although senescence is a well-established driver of aging and age-related diseases, it is a highly heterogeneous process with significant variations across organisms, tissues, and cell types. The relatively low abundance of senescent cells in healthy aged tissues poses a major challenge to the longitudinal study of senescence in specific organs, including the human lung. To overcome this limitation, we developed a positive-unlabeled learning framework to generate a comprehensive list of senescence marker genes in human lungs (termed SenSet) using the largest publicly available single-cell lung dataset, the Human Lung Cell Atlas (HLCA). We validated SenSet in a highly complex ex vivo human 3D lung tissue culture model subjected to the senescence inducers bleomycin, doxorubicin, or irradiation, and established its sensitivity and accuracy in characterizing senescence. Using SenSet, we identified and validated cell-type specific senescence signatures in distinct lung cell populations upon aging and environmental exposure.

bioRxiv (preprint)

Foundation models improve perturbation response prediction

Elijah Cole, Geert-Jan Huizing, ..., Euxhen Hasanaj, ..., Ziv Bar-Joseph, Eric P. Xing

abstract

Predicting cellular responses to genetic or chemical perturbations has been a long-standing goal in biology. Recent applications of foundation models to this task have yielded contradictory results regarding their superiority over simple baselines. We conducted an extensive analysis of over 600 different models across various prediction tasks and evaluation metrics, demonstrating that while some foundation models fail to outperform simple baselines, others significantly improve predictions for both genetic and chemical perturbations. Furthermore, we developed and evaluated methods for integrating multiple foundation models for perturbation prediction. Our results show that with sufficient data, these models approach fundamental performance limits, confirming that foundation models can improve cellular response simulations.

2025

Journal of Clinical Medicine

Toward Artificial Intelligence in Oncology and Cardiology: A Narrative Review of Systems, Challenges, and Opportunities

Visar Vela, Ali Yasin Sonay, ..., Euxhen Hasanaj, ..., Taulant Muka, Omer Dzemali

abstract

Background: Artificial intelligence (AI), the overarching field that includes machine learning (ML) and its subfield deep learning (DL), is rapidly transforming clinical research by enabling the analysis of high-dimensional data and automating the output of diagnostic and prognostic tests. As clinical trials become increasingly complex and costly, ML-based approaches (especially DL for image and signal data) offer promising solutions, although they require new approaches in clinical education. Objective: Explore current and emerging AI applications in oncology and cardiology, highlight real-world use cases, and discuss the challenges and future directions for responsible AI adoption. Methods: This narrative review summarizes various aspects of AI technology in clinical research, exploring its promise, use cases, and its limitations. The review was based on a literature search in PubMed covering publications from 2019 to 2025. Search terms included "artificial intelligence", "machine learning", "deep learning", "oncology", "cardiology", "digital twin", and "AI-ECG". Preference was given to studies presenting validated or clinically applicable AI tools, while non-English articles, conference abstracts, and gray literature were excluded. Results: AI demonstrates significant potential in improving diagnostic accuracy, facilitating biomarker discovery, and detecting disease at an early stage. In clinical trials, AI improves patient stratification, site selection, and virtual simulations via digital twins. However, there are still challenges in harmonizing data, validating models, cross-disciplinary training, ensuring fairness, explainability, as well as the robustness of gold standards to which AI models are built. Conclusions: The integration of AI in clinical research can enhance efficiency, reduce costs, and facilitate clinical research as well as lead the way towards personalized medicine. Realizing this potential requires robust validation frameworks, transparent model interpretability, and collaborative efforts among clinicians, data scientists, and regulators. Interoperable data systems and cross-disciplinary education will be critical to enabling the integration of scalable, ethical, and trustworthy AI into healthcare.

bioRxiv (accepted at Bioinformatics)

Finetuning foundation models for temporal clinical transcriptomics data

Sachin Mathur, Alexander Kagan, Peyman Passban, Hamid Mattoo, Euxhen Hasanaj, Ziv Bar-Joseph

abstract

Background: Timeseries clinical transcriptomic datasets offer the opportunity to gain insights into the dynamics of disease mechanisms/treatment responses. However, their utility in uncovering temporal patterns is often limited by high noise levels and small sample sizes. Leveraging foundational gene embeddings and incorporating interaction information can help address these challenges, improve gene network analysis, and enable the detection of subtle changes that drive disease progression or drug response. Results: We finetuned gene embeddings from foundation models using healthy tissue gene expression data and used them in temporal GNNs to model gene expression of responder and non-responders to treatment in 3 disease datasets: ulcerative colitis, Crohn's disease and psoriasis. Application of our method to these datasets confirmed known mechanisms associated with drug action, and also identified key differences between activated and repressed pathways for responders and non responders including B-Cell activation and mitochondria related activity in ulcerative colitis patients. Conclusion: Finetuning gene embeddings from foundation models provide a richer context to model gene expression data compared to using them in their naive state. Even with smaller sample sizes, results from GNN-based temporal models outperform traditional methods by detecting known mechanisms of response and unraveling role of genes and mechanisms not known to be associated with response and non-response.

ISMB 2025

Recovering time-varying networks from single-cell data

Euxhen Hasanaj, Barnabás Póczos, Ziv Bar-Joseph

abstract

Gene regulation is a dynamic process that underlies all aspects of human development, disease response, and other key biological processes. The reconstruction of temporal gene regulatory networks has conventionally relied on regression analysis, graphical models, or other types of relevance networks. With the large increase in time series single-cell data, new approaches are needed to address the unique scale and nature of this data. Here, we develop a deep neural network, Marlene, to infer dynamic graphs from time series single-cell gene expression data. Marlene constructs directed gene networks using a self-attention mechanism where the weights evolve over time using recurrent units. By employing meta learning, the model is able to recover accurate temporal networks even for rare cell types, and can identify gene interactions relevant to specific biological responses, including COVID-19 immune response, fibrosis, and aging.

ICML Workshops 2025

Multimodal benchmarking of foundation model representations for cellular perturbation response prediction

Euxhen Hasanaj, Elijah Cole, Shahin Mohammadi, Sohan Addagudi, Xingyi Zhang, Le Song, Eric P. Xing

abstract

The decreasing cost of single-cell RNA sequencing (scRNA-seq) has enabled the collection of massive scRNA-seq datasets, which are now being used to train transformer-based cell foundation models (FMs). One of the most promising applications of these FMs is perturbation response modeling: forecasting how cells will respond to drugs or genetic interventions. However, recent studies have shown that FM-based models often struggle to outperform simpler baselines, and there is a lack of understanding of the components driving performance. In this work, we conduct the first systematic pan-modal study of perturbation embeddings, with an emphasis on those derived from biological FMs, benchmarking their predictive accuracy and identifying the most successful representation learning strategies.

2024

Genome Biology

scDOT: optimal transport for mapping senescent cells in spatial transcriptomics

Nam D. Nguyen, Lorena Rosas, Timur Khaliullin, Peiran Jiang, Euxhen Hasanaj, et al., Ziv Bar-Joseph

abstract

The low resolution of spatial transcriptomics data necessitates additional information for optimal use. We developed scDOT, which combines spatial transcriptomics and single cell RNA sequencing to improve the ability to reconstruct single cell resolved spatial maps and identify senescent cells. scDOT integrates optimal transport and expression deconvolution to learn non-linear couplings between cells and spots and to infer cell placements. Application of scDOT to lung spatial transcriptomics data improves on prior methods and allows the identification of the spatial organization of senescent cells, their neighboring cells, and novel genes involved in cell-cell interactions that may be driving senescence.

Bioinformatics (ISMB Proceedings)

Integrating patients in time series clinical transcriptomics data

Euxhen Hasanaj, Sachin Mathur, Ziv Bar-Joseph

abstract

Analysis of time series transcriptomics data from clinical trials is challenging. Such studies usually profile very few time points from several individuals with varying response patterns and dynamics. Current methods for these datasets are mainly based on linear, global orderings using visit times which do not account for the varying response rates and subgroups within a patient cohort. We developed a new method, Truffle, that utilizes multi-commodity flow algorithms for trajectory inference in large scale clinical studies. Recovered trajectories satisfy individual-based timing restrictions while integrating data from multiple patients. Testing the method on multiple drug datasets demonstrated an improved performance compared to prior approaches, while identifying novel disease subtypes that correspond to heterogeneous patient response patterns.

2023

NeurIPS 2022 Competition Track (PMLR)

AutoML Decathlon: Diverse Tasks, Modern Methods, and Efficiency at Scale

Nicholas Roberts, Samuel Guo, ..., Euxhen Hasanaj, et al.

abstract

We present the AutoML Decathlon, a benchmark suite for automated machine learning spanning ten diverse application domains and modern method families, designed to stress-test AutoML systems for efficiency at scale. The competition surfaces which modern optimization and search strategies generalize across tasks that differ substantially in data modality and objective.

2022

Nature Aging

NIH SenNet Consortium to Map Senescent Cells throughout the Human Lifespan to Understand Physiological Health

SenNet Consortium

abstract

Cells respond to many stressors by senescing, acquiring stable growth arrest, morphologic and metabolic changes, and a proinflammatory senescence-associated secretory phenotype. The heterogeneity of senescent cells (SnCs) is vast, yet ill characterized. SnCs have diverse roles in health and disease and are therapeutically targetable, making characterization of SnCs and their detection a priority. The Cellular Senescence Network (SenNet), a National Institutes of Health Common Fund initiative, was established to address this need. The goal of SenNet is to map SnCs across the human lifespan to advance diagnostic and therapeutic approaches to improve human health. This Perspective lays out the impetus, goals, approaches, and products of SenNet.

Cell Reports Methods

Multiset multicover methods for discriminative marker selection

Euxhen Hasanaj, Amir Alavi, Anupam Gupta, Barnabás Póczos, Ziv Bar-Joseph

abstract

Markers are increasingly being used for several high throughput data analysis and experimental design tasks, from assigning cell types in scRNA-seq studies to selecting marker proteins in single cell spatial proteomics studies. Most marker selection methods focus on differential expression analysis, which works well for data with a few non-overlapping marker sets but is not appropriate for large atlas-size datasets. To address this, we define the phenotype cover (PC) problem for marker selection and present algorithms that can improve the discriminative power of marker sets. Analysis on several marker selection tasks suggests that these methods can lead to solutions that accurately distinguish different phenotypes in the data.

Nature Communications

Interactive single-cell data analysis using Cellar

Euxhen Hasanaj, Jingtao Wang, Arjun Sarathi, Jun Ding, Ziv Bar-Joseph

abstract

Cell type assignment is a major challenge for all types of high throughput single cell data. In many cases such assignment requires the repeated manual use of external and complementary data sources. To improve the ability to uniformly assign cell types across large consortia, platforms and modalities we developed Cellar, a software tool that provides interactive support to all the different steps involved in the assignment and dataset comparison process. We demonstrate the advantages of Cellar by using it to annotate several HuBMAP datasets from multi-omics single-cell sequencing and spatial proteomics studies. Cellar is open-source and includes several annotated HuBMAP datasets.