Datasets and tools

We develop and release research software, genome assemblies and annotations, population-genomic datasets, breeding and mapping populations, phenotyping resources and reproducibility workflows for crop evolutionary genomics and breeding.

Where possible, sequence data are deposited in ENA/NCBI, released variant datasets in EVA, phenotype and imaging datasets in public repositories, and research software in GitHub and/or Zenodo.

Different links on this page refer to different levels of the underlying research resource:

  • a BioProject or ENA study normally contains the underlying raw sequence reads;
  • an assembly accession identifies a particular genome assembly;
  • EVA provides released and reusable variant datasets;
  • Dataverse/Figshare/Zenodo records may contain phenotypes, variants, assemblies, annotations or software;
  • supplementary tables frequently contain accession metadata, phenotype matrices, pedigrees, marker genotypes and mapping results that are not deposited elsewhere;
  • GitHub repositories may contain either reusable tools or publication-specific reproducibility code.

For reuse, please cite the associated publication and, where appropriate, the sequence accession, dataset DOI or software release.


Quick reference of papers, datasets and code — and where to find it

This section is deliberately redundant. It provides the fastest route from a publication to the corresponding sequence data, code, phenotypes, supplementary tables and other reusable resources.


Tools and reusable workflows

LegumeDiscovery

LegumeDiscovery provides an entry point for discovering and navigating curated legume breeding datasets and analysis-ready resources generated through our work on crop improvement.

The aim is to make breeding data easier to find, interpret and reuse, connecting germplasm, experiments, environments and traits with harmonised phenotype datasets and downstream analyses.

Activities around these datasets include:

  • curation and harmonisation of breeding-trial data;
  • consistent trait definitions and metadata;
  • integration of multi-site and multi-season experiments;
  • generation of analysis-ready phenotype matrices;
  • genotype quality control and imputation;
  • quantitative-genetic analyses;
  • genomic prediction and association analyses.

Resource: DeVegaGroup on GitHub


autoBLUEs/BLUPs

autoBLUEs/BLUPs automates mixed-model analysis of replicated and multi-environment phenotyping experiments to generate genotype-level BLUEs — best linear unbiased estimates — and BLUPs — best linear unbiased predictions.

The workflow provides a reproducible bridge between raw experimental observations and downstream analyses including:

  • GWAS;
  • genomic prediction;
  • genotype ranking and selection;
  • multi-environment trial analysis;
  • estimation of genetic and environmental effects;
  • genotype × environment analysis;
  • comparison of traits across experiments, seasons and locations.

Resource: DeVegaGroup on GitHub


SeedAnalyser

SeedAnalyser is our computer-vision and machine-learning framework for extracting quantitative information from scanned seeds and classifying seed species.

Code: SeedClassifier

The repository includes:

  • ImageJ macros;
  • OpenCV-based segmentation;
  • Cellpose-based segmentation;
  • extraction of seed dimensions and image-derived features;
  • comparison of conventional image-analysis pipelines;
  • classical machine-learning classifiers;
  • deep-learning / ResNet classification;
  • trained models;
  • model evaluation;
  • open-set classification experiments.

Paper: Integrating machine learning, deep learning, and image analysis for seed species classification.


Relative Averaged Alignment and Relative Coverage

We developed two alignment-based approaches for identifying ancestry changes and introgressed chromosome segments:

  • Relative Averaged Alignment (RAA)
  • Relative Coverage (RC)

Code: introgressions_by_relative_depth

Paper-specific analyses: Structural-diversity-in-banana-cultivars

Paper: Characterizing subgenome recombination and chromosomal imbalances in banana varietal lineages.

Sequence data: ENA PRJEB62882.


Basecall2Assembly

Basecall2Assembly is a Snakemake workflow for moving from raw Oxford Nanopore Technologies sequence data towards a genome assembly.


QCPipeline

QCPipeline provides a reproducible workflow for quality control and evaluation of genome assemblies.

Repository: QCPipeline

It complements Basecall2Assembly and the project-specific assembly workflows distributed with our genome publications.


Genome assemblies and annotations

This section contains reference genomes, chromosome-scale assemblies, draft genomes, structural and functional annotations, and physical genome resources.

Population-level genetic diversity datasets derived from these or related studies are listed independently below.


Urochloa humidicola cv. Tully — haplotype-complete octoploid genome

A haplotype-resolved chromosome-scale reference for the octoploid tropical forage grass Urochloa humidicola cv. Tully.

The analysis repository contains workflows for long-read assembly, scaffolding, read mapping, assembly assessment, BUSCO, KAT, dot plots, orthology, phylogenomics and analysis of chromosome/subgenome composition.


Urochloa decumbens cv. Basilisk — haplotype-resolved allotetraploid genome

A chromosome-level haplotype-resolved assembly of the apomictic allotetraploid Urochloa decumbens cv. Basilisk.

The repository contains workflows for assembly, ancestry clustering, BUSCO, Merqury, repeat analysis, genome-composition plots, dot plots and synteny.


Urochloa ruziziensis CIAT 26162 — reference genome and annotation

A reference genome for diploid Urochloa ruziziensis CIAT 26162 developed to support comparative genomics and mapping of natural variation in tropical forage grasses.

The associated BRX 44-02 × CIAT 606 family is listed separately in the population section.


Miscanthus sinensis DH1 — chromosome-scale reference genome

A chromosome-scale reference genome for Miscanthus sinensis, assembled into its 19 chromosomes and used for comparative, evolutionary and population-genomic analyses.

A four-cross genetic map containing 4,298 uniquely assigned markers was used in validating and anchoring the chromosome-scale assembly.


Miscanthus sacchariflorus cv. Robustus 297 — draft genome and annotation

A draft genome resource for the bioenergy grass Miscanthus sacchariflorus cv. Robustus 297.

The Zenodo archive includes genome FASTA, chromosome-anchored sequence, GFF3 gene annotation, functional annotation and AGP anchoring information.


Sugarcane hybrid CC 01-1940 — chromosome-level genome and annotation

A chromosome-level genome assembly of the high-yielding Colombian sugarcane hybrid CC 01-1940.


Lolium perenne P226/135/16 — BAC physical map and genome-sequence resource

A physical genome and BAC-sequence resource connecting the Lolium perenne genome with genetic mapping and association analyses.

Important supplementary resources include marker sequences, SNP positions, BAC-library statistics, complete FPC/LTC physical maps and genome-coverage statistics.

The associated 716-genotype European GWAS panel is listed separately below.


Red clover (Trifolium pratense) — draft genome and annotation

A draft reference genome and gene annotation for red clover.


Genetic diversity, association and mapping populations

This section contains crop diversity panels, germplasm collections, GWAS populations, breeding populations and segregating families.

The ordering is intended to make the larger crop-population resources easy to find first.


Vietnamese native-rice diversity panel

A whole-genome resequencing resource comprising 672 Vietnamese native rice accessions, developed for analysis of crop diversity, population structure, trait association and genomic regions affected by breeding.

Data

Particularly reusable supplementary tables

  • Table S1: accession identities, National Genebank numbers, local names, collection locations, sequencing/mapping statistics and population assignments.
  • Table S2: comparative dataset of 3,635 rice varieties.
  • Table S3: phenotypic measurements for 20 traits across the 672 accessions.
  • Table S4: phenotype definitions and abbreviations.
  • Table S5: phenotype summary statistics and population comparisons.
  • Table S6: nucleotide diversity by subpopulation.
  • Table S7: GWAS results and reported QTL.
  • Table S8: genes associated with QTL.
  • Table S9: IRRI accessions.
  • Table S10: definitions of the SNP datasets used in the analyses.
  • Table S11: functional annotation of the SNP set.

The later selection study adds Supplementary Tables S1–S13 describing selected genomic regions, population-genomic statistics and candidate genes.


Banana diversity panel — population structure, introgression and GWAS

A whole-genome resequencing panel of cultivated bananas used for studies of ancestry, subgenome recombination, chromosomal imbalance and agronomic-trait association.

Data

Papers

The same genomic panel can be reused for:

  • ancestry inference;
  • population structure;
  • introgression;
  • chromosome imbalance;
  • structural diversity;
  • association mapping.

Phenotypes covering morphology, fruit quality and yield are supplied through the GWAS paper and its supplementary datasets.


Legume populations

Red clover diversity panel

A diversity resource designed to characterise population structure, adaptation and useful genetic variation across European and Asian red clover germplasm.

Population

  • 75 accessions
  • 70 natural populations/ecotypes and five commercial varieties
  • 640 individual plants originally sampled

Data

Supplementary resources

The supplementary spreadsheets contain:

  • SNP statistics;
  • accession-level heterozygosity;
  • AMOVA;
  • outlier loci and candidate genes;
  • genotype–environment associations;
  • survival;
  • flowering;
  • plant architecture;
  • geographic and climate groupings;
  • commercial-variety metadata.

Particularly useful tables include:

  • Table S4: candidate loci showing signatures of selection.
  • Table S6: survival/mortality.
  • Table S7: vegetative-growth phenotypes.
  • Table S8: geographic/climatic grouping.
  • Table S9: commercial-variety metadata.

Common bean diversity panel — determinacy and photoperiod

A whole-genome resequencing panel used to dissect selection for growth habit, determinacy and photoperiod sensitivity.

Data and code


Liborino common-bean germplasm panel

A collection of 44 Liborino-type common bean accessions used for germplasm characterisation, adaptation, yield evaluation and participatory selection.

No separate public sequence BioProject was identified for this experiment; the primary reusable resources are the germplasm metadata and multi-site phenotype datasets published with the article.


Temperate grasses and bioenergy crops

Lolium perenne European ecotype GWAS population

A European ryegrass diversity population consisting of:

  • 716 diploid genotypes
  • from 90 accessions / collection sites
  • phenotyped over two years;
  • genotyped using a custom Lolium Infinium SNP array.

Resources

The BioProject contains the BAC/genome-sequence resource; the GWAS population genotype and phenotype information is supplied primarily through the publication and supplementary datasets.


Miscanthus diversity, admixture and mapping populations

The Miscanthus sinensis genome study also generated substantial population-genomic and mapping resources.

Data

The analyses cover M. sinensis, M. sacchariflorus, interspecific admixture and the evolutionary origin of M. × giganteus.

A four-cross genetic map with 4,298 uniquely assigned markers provides an additional family-based mapping resource.


Potato TON multi-environment breeding population

A 380-genotype tetraploid potato breeding panel from the International Potato Center, evaluated for late-blight resistance across multiple environments.

Phenotypes

Genotypes

Supplementary tables

  • Table S1: population assignments and parentage of 380 genotypes.
  • Table S2: genotype BLUEs for rAUDPC by environment.
  • Table S3: pedigree BLUPs.
  • Table S4: genomic BLUPs.

Paper: Global multi-environment resistance QTL for foliar late blight resistance in tetraploid potato with tropical adaptation.

This is a particularly reusable breeding dataset because field data, genotype calls and derived BLUE/BLUP values are all publicly available.


Tropical forage-grass populations

Guinea grass (Megathyrsus maximus) diversity and GWAS panel

A diversity panel of 124 genebank accessions used to investigate the genetic basis of agronomic, biomass and nutritional traits.

Traits include plant architecture, flowering, biomass production, protein, fibre and digestibility.


Napier grass (Cenchrus purpureus) global diversity collection and progeny

A global resequencing resource containing 450 Napier grass genotypes sampled from international germplasm and breeding collections.

The study also includes 109 open-pollinated progeny from 14 maternal genotypes.

Data

Supplementary datasets

  • Table 1: accession metadata and phenotype information.
  • Table 2: detailed trait performance.
  • Table 3: PCA results.
  • Table 4: chromosome-level SNP statistics.
  • Table 5: population/admixture assignments.
  • Table 6: interspecific hybrids.
  • Tables 7–10: marker-trait associations and candidate regions.

Urochloa 111-accession population-genomic panel

A multispecies tropical-forage panel designed to investigate how reproductive mode, hybridisation and genome composition influence population structure.

Data

The resource includes 111 genetically distinct accessions and supports analyses of:

  • species relationships;
  • genetic differentiation;
  • admixture;
  • reproductive mode;
  • subpopulation structure.

Urochloa genomic-composition and cytogenomic collection

A broad germplasm characterisation resource integrating taxonomy, genome composition, cytogenetics, repetitive-DNA analysis and whole-genome sequencing.

Data

Supplementary datasets

  • Table S1: accessions, genome composition, growth habits and geographic distributions.
  • Table S2: sequencing data for the nine WGS accessions.
  • Tables S3–S4: genome-specific candidate sequences and probes.
  • Tables S7–S10: repeat, k-mer, genome-specific sequence and transposable-element analyses.

Urochloa apomixis and aluminium-tolerance F1 mapping family

The apomixis and aluminium-tolerance studies use the same interspecific BRX 44-02 × CIAT 606 family and are therefore presented here as a single reusable genetic resource.

Population

  • approximately 169 F1 progeny
  • sexual U. ruziziensis BRX 44-02 × apomictic U. decumbens CIAT 606 cv. Basilisk

The population has been used to study both:

  • reproductive mode and the apospory-specific genomic region;
  • natural variation in aluminium tolerance.

Genome and sequence resources

Apomixis paper

A Parthenogenesis Gene Candidate and Evidence for Segmental Allopolyploidy in Apomictic Brachiaria decumbens

Reusable supplementary resources include:

  • GBS sequencing depth;
  • marker and primer information;
  • marker genotype scores for the family;
  • marker segregation classes;
  • linkage-map markers;
  • comparative positions relative to foxtail millet;
  • reproductive-mode phenotypes.

Aluminium-tolerance paper

A new genome allows the identification of genes associated with natural variation in aluminium tolerance in Brachiaria grasses

Reusable supplementary resources include:

  • aluminium-tolerance phenotypes;
  • genetic markers and association/mapping results;
  • candidate genomic regions;
  • genome structural and functional annotation.

Together, the two studies make this one of the more extensively characterised Urochloa segregating families.


Urochloa spittlebug-resistance F1 population

A breeding population of 339 interspecific F1 hybrids used to dissect resistance and tolerance to Aeneolamia varia spittlebug nymphs.

Data

The combination of genomic data and thousands of plant images makes this both a mapping-population dataset and an independently reusable computer-vision resource.


Phenotyping and image datasets

Urochloa spittlebug plant-damage image collection

A high-throughput image dataset generated from experimentally phenotyped Urochloa plants challenged with spittlebug nymphs.

The collection can be reused independently for development and benchmarking of quantitative plant-damage and computer-vision methods.


Seed image-analysis and classification resources

Resources associated with SeedAnalyser include seed images, quantitative image descriptors, trained models and model-evaluation code.

The repository supports reuse of segmentation, feature extraction, classical machine learning and deep-learning components independently of the complete pipeline.


Transcriptomic datasets

Miscanthus drought-response transcriptomics

RNA-seq data examining physiological and transcriptional responses to drought across Miscanthus material.


Miscanthus biomass, starch and sucrose transcriptomics

A transcriptomic resource linking expression of starch- and sucrose-metabolism genes with variation in biomass yield.

Supplementary resources include phenotype measurements, normalised expression values, differential-expression statistics, functional annotations and regulatory/network analyses.


Urochloa drought-response transcriptomics

A transcriptomic experiment examining drought responses in contrasting Urochloa hybrid genotypes.


Reuse and cite!