Who we are

We are the Crop Evolution & Adaptation Lab (De Vega Lab) at the Earlham Institute (Norwich Research Park, UK).

Our lab is a multidisciplinary genomics group studying how genome evolution and genetic variation shape agriculturally important traits, and how this knowledge can be translated into crop improvement. We work particularly on the challenges created by hybridisation (interspecific complexity), polyploidy, structural variation, and asexual reproduction.

Our established collaborations with breeding companies, crop research institutes, and international consortia keep the programme grounded in real agricultural problems and stakeholder needs. Increasingly, our work also involves developing genomic resources, analytical workflows, and FAIR data standards and infrastructure that allow these communities to discover, analyse, and reuse complex genomic and phenotypic data.

Our code, workflows, and datasets are usually released in accordance with FAIR principles, enabling reproducibility, reuse, interoperability, and cumulative improvement by the wider research community.

On this page

Crop Evolution & Adaptation

One of the central questions in biological research is how phenotypes arise. In this context, our research aims to understand how the evolutionary processes of hybridisation and polyploidisation shape genetic diversity across the genome and how this diversity manifests in agronomic traits that could be leveraged in breeding. We work at the interface of genome evolution, genomic variation, and crop improvement.

Our research investigates the mechanistic basis of phenotypic variation, develops approaches to predict phenotypes, and works with breeding and biotechnology partners to translate genomic discoveries into crop improvement. Our goal is not simply to catalogue diversity, but to understand and interpret variation well enough to support better biological prediction and practical breeding decisions.

Our Research objectives

Our first objective is to develop and apply genomic approaches that explain the mechanistic manifestations and evolutionary consequences of hybridisation, polyploidisation, introgression, and asexual reproduction. In practice, this means resolving and interpreting gene flow across lineages, structural and copy-number variants, dosage effects, expression dominance, homoeologous exchange, and other forms of complex genomic variation.

Because these same processes are extensively exploited or encountered in crop improvement, our second objective is to translate knowledge of genome evolution and genetic variation into approaches, resources, and predictions that improve breeding. This provides a direct pathway from fundamental genome biology to strategies that accelerate genetic gain and support sustainable food production.

A third and increasingly important objective is to make genomic and phenotypic information more usable by research and breeding communities. We therefore contribute to open workflows, shared genomic resources, standards, ontologies, and community-facing data infrastructure that support reproducible analysis and reuse across projects and institutions.

Biological systems

Many major crops are polyploid or of hybrid origin, particularly among cereals, tubers, and fibre crops. We focus on particular allopolyploid systems, especially bananas and several feed crops, and on species with multiple domestication gene pools, including rice and beans. Hybridisation is a common mechanism of diversification and, in plants, is often associated with polyploidisation. Chromosomal duplication can buffer against meiotic irregularities and regulatory mismatches more effectively than a diploid background, but can also introduce its own challenges, such as copy-number variants and dosage effects.

Hybridisation and polyploidisation can generate extensive structural, copy-number, regulatory, and epigenetic variation. Using comparative, population and pangenomic approaches, we quantify these forms of variation and investigate how they persist, recombine, and affect traits. Asexual reproduction is also relevant in this context, because clonal propagation, apomixis, parthenogenesis, and related reproductive modes can alter how genetic variation is generated, fixed, and maintained.

How we work

We combine genome assembly, population genomics, genome-wide association analysis, pangenomics, long-read sequencing, quantitative genetics, high-throughput phenotyping, and phenotype prediction. Our computational work spans from raw sequencing data through genome assembly and annotation, genetic variation analysis, multi-environment trial analysis, and predictive modelling.

We are especially interested in forms of variation that are often poorly represented by single-reference or SNP-centred analyses, including structural variation, copy-number variation, dosage effects, homoeologous exchange, introgression, and epigenetic variation.

Complex and polyploid crops provide useful systems because they concentrate many of the challenges now facing genomics more generally: multiple haplotypes, introgression, structural change, variable dosage, incomplete representation by single reference genomes, and rapidly increasing genomic data volumes.

We aim to develop approaches that are biologically informative, computationally reproducible, and useful under real breeding conditions. This includes releasing reusable workflows and analytical code, contributing genomic resources, and working with collaborators to establish data standards and FAIR practices.

From genomic variation to community resources

An increasing part of our work concerns the transition from analysing genomic datasets within individual studies to developing resources that allow wider communities to discover, integrate, analyse, and reuse them.

Within the Horizon Europe Legume Generation programme, we co-lead development of the project’s digital Knowledge Centre Legume Discovery. The resource is being developed as a unified environment for phenotypic, multi-environment trial, and genotyping data generated across several crop Innovation Communities.

Our group contributes scientific leadership and computational development to the platform. Current work includes FAIR data management, data and metadata harmonisation, trait ontologies, integration of datasets from multiple project partners, and analytical functionality for multi-environment trials.

A dedicated module compares mixed-linear models and generates BLUEs and BLUPs across locations and seasons, while backend development is extending the platform towards marker-data management and downstream GWAS and marker-to-target discovery.

This work reflects our broader view that useful genomic infrastructure should combine high-quality data, reproducible analytical methods, interoperable standards, and sustained engagement with the communities that generate and use the data.

Future direction

A longer-term ambition of the lab is to develop predictive genomics approaches in which biological behaviour becomes increasingly predictable from genome, phenotype, and environment, and in which models can improve as new trials, genomic resources, and experimental evidence become available.

We use machine learning and AI where they provide a demonstrable advantage. Our work already includes benchmarking machine-learning and deep-learning approaches for classification and phenotype prediction, alongside mixed and statistical models. We are particularly interested in combining these methods with biologically interpretable genomic features rather than treating prediction as an end in itself.

Our longer-term workflow is therefore one of discovery, integration, prediction, and validation: identifying biologically meaningful genomic variation; organising and exposing relevant genomic and phenotypic information through reusable resources; developing predictive models; and testing those predictions through trials, bioassays, or collaborator-led experimental validation.

Acknowledgement

We are supported by the Earlham Institute’s research office and other support teams, the Transformative Genomics group at the Earlham Institute, BDE and support teams, the Norwich Bioscience Institute’s research computing, contracts, and finance teams, horticultural services at the John Innes Centre, and genebank personnel from around the world.

We are funded by the BBSRC, part of UK Research and Innovation (UKRI), through the Earlham Institute Strategic Programme Grant Decoding Biodiversity BBX011089/1 and its constituent work package BBS/E/ER/230002B (Decode WP2); the Institute Strategic Programme Grant Resilient Crops BB/X011062/1 under project BBS/E/ER/230004A; the BBSRC-funded Norwich Research Park Biosciences Doctoral Training Partnership grant BB/T008717/1; the NERC-funded ARIES Doctoral Training Landscape Grant; and the Innovate UK-funded grants 10077978 (Legume Generation) and 10102570 (KTN).

On this page