RESEARCH

SEPTEMBER 11, 2026

Rare Cancer, Small Data: How Do You Build a Model With 1,000 Patients a Year?

One patient's extraordinarily deep data, and what dogs and AI can add to it

One patient's extraordinarily deep data, and what dogs and AI can add to it

By

Clyra Bio

In November 2022, Sid Sijbrandij, co-founder of GitLab, was diagnosed with osteosarcoma. The tumor in his spine was an unusual presentation of an already rare cancer. From that point on, Sid approached the disease in “founder mode”: rather than relying only on the standard treatment pathway, he took an unusually proactive, data-driven approach to understanding how his cancer was changing.

Between December 2022 and early 2025, Sid’s tumor was sampled four times. Each sample was analyzed in molecular depth using a combination of technologies. Together, these measurements created a longitudinal view of how the tumor and its surrounding microenvironment evolved over the course of treatment.

However, most patients with osteosarcoma do not have access to this level of monitoring. If molecular profiling is performed at all, it is often done at only a single timepoint, typically around diagnosis or after initial chemotherapy. As a result, the picture we usually have of a patient’s cancer is only a snapshot. For much of the time between scans or recurrences, we have limited visibility into how the tumor is changing at the molecular level. Patients and clinicians often have to make decisions without knowing whether the biology of the disease has shifted until those changes become clinically apparent.

Fortunately, Sid chose to make his data publicly available. Rich longitudinal datasets like his can provide valuable clues for developing more generalizable and cost-effective ways to understand and monitor cancer in greater biological details. By studying tumors that have been profiled repeatedly over time, we can begin to identify which features in an early sample are associated with what happens later, such as how aggressive the disease may become, how the immune system interacts with the tumor, and how the tumor may respond to different treatments.

This is also central to what we are building at Clyra. Our goal is to learn from deeply characterized cases like Sid’s and develop models that can extract more predictive value from a patient’s initial molecular profile. Ultimately, we hope this can help patients and clinicians choose treatments more effectively and monitor disease with fewer, less resource-intensive measurements.

Different tumor analyses at different timepoints

Sid’s tumor samples were collected at four different timepoints and analyzed using several molecular techniques, including whole-genome sequencing (WGS), RNA sequencing (RNA-seq), and other genomic analyses. Together, these analyses provided a longitudinal view of how his tumor changed over time at the molecular level. Timepoint T0 was taken at diagnosis in December of 2022, T1 in June 2024 at the first recurrence of his cancer, T2 in January 2025 at the second recurrence, and T3 in April of 2025 at the time of the surgical excision of his recurred tumor after he received an experimental fibroblast-activation protein (FAP) radioligand therapy in Germany.

What Sid’s Tumor Revealed Over Time

The tumor and its environment changed. Tumor cells exist among other cells, including immune cells, blood vessels, and fibroblasts. The surrounding population of cells is called the tumor microenvironment (TME). When a tumor biopsy is collected, both tumor cells and surrounding TME components are sampled. Across Sid’s samples, the tumor cells make up less and less of what was collected in the samples over time, while immune cells make up more. At the final time point, of the 8,717 cells profiled, the three largest populations were different types of white blood cells: neutrophils (43%), T-cells (16%), and macrophages (14%). We see the opposite trend in the recurrent and metastatic tumors in the public cohort of 11 osteosarcoma patients that we compared to Sid’s final time point. In these patients, tumor cells make up the majority of the samples.

Sid’s cells move out of tumor territory on a map of other patients’ cells. To compare tumor samples across batches and time points, we use Universal Cell Embedding (UCE), a single-cell foundation model that maps cells into a shared biological representation space. UCE uses each cell’s gene-expression profile to generate a cell-level embedding that captures its underlying biological state. We then project these embeddings into a shared space to visualize cell populations across time points.

When Sid’s single cells are mapped into the same UCE space, we observe a longitudinal shift in the tumor microenvironment: cell populations occupying tumor and tumor-associated regions decrease over time, while populations corresponding to T-cells and neutrophils increase. This suggests a transition from a tumor-dominant microenvironment toward a more immune-enriched state.

Animated UCE cell-type trajectory showing Sid's longitudinal tumor microenvironment from T0 through T3

His disease looks progressively less like advanced disease. We can also use the UCE map to ask which fraction of Sid’s cells have their closest matches among other patients’ tumor cell populations that had already recurred or spread. Across the three timepoints (T1-T3) with measured single-cell data, that fraction falls from about 15 percent to under 9 percent.

Animated UCE trajectory comparing Sid's cells with primary, recurrent, and metastatic osteosarcoma reference populations

The proportion of cells expressing drug targets moved. Several of the newer drugs being developed for osteosarcoma aim at a specific protein on or around the tumor cells. B7-H3, also called CD276, is one of the most prominent protein targets. The spatial data shows where that protein is actually being made, and that production changes over the four timepoints. B7-H3 covered about a third of the diagnostic tissue (T0), rose to over half at the first recurrence (T1), and by the final timepoint (T3) shrank again to less than one sixth of cells. Visually, at T2, the expression of B7-H3 looks nearly continuous across the tissue while expression seems to be scattered in smaller clusters at T3. Expression patterns of EGFR and FAP, two other candidate targets, look similar to CD276 across the timepoints. These images produced over time illustrate how a drug aimed at a target that was abundant at diagnosis may be aiming at something that is no longer there at recurrence.

Xenium spatial transcriptomics sections showing CD276, EGFR, and FAP expression across six tumor samples over time

Caption: Xenium spatial sections from Sid’s tumor. Rows are genes, including CD276 (B7-H3), EGFR, and KERA. Columns are tissue sections in time order: two sections from diagnosis (T0-B3, T0-C3), one each from the first and second recurrences (T1, T2), and two from the final surgical excision (T3-A, T3-B). Grey background is tissue morphology, while colored dots are cell centroids where the gene was detected. The inset percentage is the fraction of segmented cells with at least one transcript of that gene.

The genome tells a consistent story. At each timepoint, the WGS data were analyzed using three independent DNA variant-calling pipelines. To reduce the risk of false positives, we report only mutations identified by at least two of the three pipelines. Tumor mutational burden stayed low and the tumor stayed microsatellite stable throughout, which is exactly what the osteosarcoma literature would predict for a cancer driven by structural and copy-number changes rather than by point mutations. Amplification of MYC, which published cohorts associate with three-year survival of about 34 percent versus 70 percent without it, was present at diagnosis (T0) and not detected at any later timepoint. It’s possible that the tumor cells with this amplification were successfully destroyed during treatment which may have improved Sid’s prognosis, but tumor purity in the later samples was much lower (as low as 16 percent), so it is possible that the mutation was still present but went undetected in these later samples.

The finding that changed Sid’s later treatment was an expression finding. Standard genomic panels had not identified any actionable mutations in Sid’s tumor. One key alteration was MYC amplification, a large copy-number change rather than a recurrent point mutation that can be readily matched to an existing targeted therapy. On the other hand, single-cell RNA sequencing showed something the panels could not: his tumor cells were producing fibroblast activation protein (FAP). FAP is normally made by the fibroblasts that build a tumor’s structural scaffolding instead of by tumor cells themselves. That made him a candidate for a FAP-targeted radioligand, and a 68Ga-FAP PET scan in October 2024 confirmed his tumor took up a FAP-binding tracer. The radioligand therapy he received in Germany followed from that scan.

A great resource that doesn’t scale

Sid’s data is one person’s data. It presents an inspiring case study of a patient sampled at four critical timepoints over two years with each sample analyzed in depth, revealing tremendously important details that guided his enrollment in an emerging therapy. The data also shows how pivotal findings for Sid’s care were not detectable at all timepoints or via all testing modalities.

Whole-genome sequencing, single-cell RNA sequencing, and spatial transcriptomics are all well-established methods, and the cost of each continues to decrease as technologies improve. Still, the use of more than one of these methods per patient at multiple timepoints during diagnosis and treatment is far from routine. The total monetary, time, and logistical costs involved with serial sampling, laboratory processing, computational work, and physician interpretation is enormous.

If that weren’t enough, three physical barriers work against osteosarcoma specifically: decalcification of bone specimens damages nucleic acids; post-chemotherapy resection specimens often fail tumor-purity thresholds (because good response means little viable tumor left); and the largest routine-sequencing study of children with cancer under-sampled bone tumors because those patients are biopsied at separate regional sarcoma units. The net effect of these barriers is that the diagnostic biopsy is often the only specimen good enough to sequence. This makes Sid’s four timepoints extraordinary in addition to being cost- and resource-intensive.

Beyond the difficulty of monitoring disease progression frequently across multiple data modalities, osteosarcoma is also challenging to study because it is such a rare cancer. Around 1,000 people are diagnosed each year in the United States, about half of them children and adolescents. Cohorts are small and trials are slow to reach the numbers needed for adequate statistical power.

Interestingly, osteosarcoma is not a rare cancer in dogs. Canine osteosarcoma is about ten times more common than the human form, and both species share similar genetic changes, patterns of metastasis, and responses to treatment. The canine disease progresses more quickly, which means that the clinical events that we tend to see over years with human patients (e.g. diagnosis, remission, relapse) happen over months in the dog. This gives us an opportunity to think differently about a rare human cancer: by learning from canine patients, we may be able to generate insights more quickly and use them to better understand and predict human osteosarcoma.

From Intensive Monitoring to Scalable Risk Prediction

Sid’s data shows the value of frequent molecular monitoring for understanding how a patient’s cancer changes over time, but this level of testing is expensive and difficult to reproduce for most patients. To make these insights more accessible, we need prediction models that can use limited measurements to estimate survival outcomes and track changes in disease more cost-effectively. This challenge became one of the motivations behind starting Clyra.

For osteosarcoma, however, building reliable prediction models is especially challenging because the disease is rare and human patient datasets are relatively small. To overcome this limitation, we borrow information from canine osteosarcoma, which shares important biological features with the human disease. By combining insights from dogs and humans, we aim to strengthen our models and make more accurate predictions even when human data are limited.

The idea of learning about human cancer from dogs that develop it naturally is not new. The Comparative Oncology Trials Consortium at the National Cancer Institute has spent two decades running clinical trials in pet dogs with spontaneous cancers and building the evidence that their tumors can inform human drug development, with osteosarcoma central to that work throughout. What has changed recently is not the premise but the measurements available to test it. Abundant sequencing data are now available in both species, and new AI foundation models make it possible to analyze them within a shared biological framework.

We are not the first to notice this. Patkar and colleagues, writing in Clinical Cancer Research in 2024, built a classifier on tumor microenvironment subtypes from 245 canine tumors and applied it to human osteosarcoma datasets, where it stratified progression-free survival, five-year metastasis, and response to chemotherapy. Based on these findings, we built a survival prediction model using a feedforward neural network trained on both human and canine data. Compared with a model trained on human data alone, the cross-species model showed a striking improvement in predictive performance, suggesting that canine data can provide valuable additional information for predicting outcomes in human osteosarcoma.

To evaluate predictive performance, we used Harrell’s concordance index (c-index). A c-index of 0.5 indicates performance no better than random chance, while a c-index of 1.0 represents perfect prediction. Trained on human osteosarcoma data alone, our model scored 0.545 on held-out human patients, getting 72 of 132 patient pairs in the right order. That is close to guessing, and this is not surprising given the rarity of the cancer and the resulting scarcity of data.

Adding canine osteosarcoma to the training data raised the c-index to 0.682, which is 90 correct out of the same 132 pairs. Twenty-five pairs that the human-only model had ordered backwards came out right, and 7 that had been right came out backwards, for a net gain of 18. The improvement appeared separately in all three human cohorts we tested, and it held under bootstrap resampling, though across a wide range of plausible values.

Bar chart comparing Harrell's C-index for human-only and human-plus-canine osteosarcoma survival models

Caption: Harrell’s C-index used to evaluate two survival models: human-only training (gray) vs. human + canine training (blue). Models are evaluated on held-out human cohorts, while canine data used in training only. The dashed line at 0.5 marks chance-level prediction. Adding canine data improved concordance in every cohort, raising pooled performance from 0.545 to 0.682, and rescued two cohorts where the human-only model performed at or below chance.

Our model has a wide range of potential applications. One example is to estimate how a patient’s disease is likely to behave, so that monitoring and treatment timing can be matched to risk rather than to a generic schedule. With additional drug-specific data, we can fine-tune the model to predict treatment response. This could help identify which patients are most likely to benefit from a given therapy, while helping pharmaceutical companies better understand why treatment responses vary across patients.

Breaking the Cross-Species Translation Barrier

One of the biggest challenges in using canine data to inform human osteosarcoma is translating biological information across species. Traditionally, comparative genomics has relied heavily on predefined gene annotations and known orthologs, which limits how much information can be transferred from dogs to humans.

Human osteosarcoma single-cell UCE embedding with labeled immune, stromal, endothelial, osteoclast, and tumor cell populationsCanine osteosarcoma single-cell UCE embedding with labeled immune, stromal, endothelial, osteoclast, and tumor cell populations

Caption: Single-cell osteosarcoma data from human (left) and canine (right) tumors were embedded using UCE and integrated with Harmony. Despite coming from different species, the two maps show a remarkably similar structure: cytotoxic CD8 T cells occupy the upper region, B and plasma cells form a nearby cluster, endothelial cells appear on the left, matrix CAFs in the lower-left, osteoclasts in the lower-center, and C1QC macrophages near the center. In other words, cells with similar biological identities occupy similar neighborhoods in the shared embedding space, regardless of whether they came from a human or a dog.

AI foundation models offer a new way to address this problem. By embedding genes, cells, or other molecular features from different species into a shared representation space, we can identify biological similarities based on their proximity in that space rather than relying only on existing annotations. This creates a more annotation-agnostic approach to cross-species modeling.

We believe this could substantially expand the amount of canine information that can be incorporated into human prediction models, improving statistical power and helping uncover biological signals that would be difficult to detect from rare human osteosarcoma cohorts alone.

From one patient to a more scalable future

Sid’s case shows what becomes possible when we can observe cancer deeply and repeatedly over time. But the goal should not be for every patient to require four biopsies, multiple sequencing technologies, and years of intensive molecular monitoring to gain those insights.

Instead, we want to learn from deeply characterized cases like Sid’s and use those lessons to build models that can extract more information from fewer measurements. For a rare cancer like osteosarcoma, canine patients offer an important opportunity to expand the amount of biological and clinical information available. New AI models give us a way to bring that information into the same biological framework and ask what can truly translate across species. Rare cancer does not have to mean learning from rare data alone.

Our hope is that, over time, this approach can make personalized cancer monitoring and treatment prediction more accessible. Not only for exceptional cases like Sid’s, but for many more patients facing rare cancers.