DGX Spark · GB10 · Parabricks WGS


Executive summary

Control-calibrated result

Clone groups sit at (or just above) the same-tree control's noise floor, while every genuinely distinct cultivar is ·× higher. Names in the collection don't track genotypes: many are synonyms.

How to read the numbers — three metrics

Clonality is measured three ways, and they are not interchangeable: absolute counts differ ~1000× between them. Only the ranges within a card compare. The heatmap below uses the middle metric; the biological verdict rests on the right-hand one.

Relatedness map

Pairwise fixed differences (sites where both samples are homozygous for different alleles — the clonality metric, immune to heterozygous-call noise). Cell shade = similarity (darker = closer); colored outlines + labels mark clone groups. Numbers shown for clone/near-clone pairs; hover any cell for exact counts. Ordered by UPGMA on the same metric. Note: the counts here still carry a residual mapping-artifact floor — · even between two leaves off one tree — and the artifact-corrected difference for true clones is ~0; see "Under review" below.

The clone groups

Each group is one genotype under several names. Within-group fixed differences run · per Mb of comparable sites — at or near the same-tree control's · — against · per Mb between distinct cultivars. · groups carry an under-review flag where the genome conflicts with the horticultural record.

Kinship & relatedness — accepted-methods cross-check

IBD sharing decomposition (Z0/Z1/Z2)

True relationship degrees

The honest limit

Independent marker classes & genome-wide screens

Every clone call above rests on SNPs. bcftools and DeepVariant agree on them, which controls for the caller — but both read the same alignments and score the same variant class. These checks change the marker class or look outside the reference entirely, so a systematic problem in the SNP pipeline would not survive all of them. Each one carries the control that makes its verdict mean something, and the negatives are reported as fully as the positives.

Structural variants — the strongest independent class

What the color and stripe phenotypes are not

Non-host DNA — real, and not the answer

Inbreeding, and a free positive control

Somatic sport rate — a NULL its own control limits

A second reference genome — measured, not adopted

Provenance cross-check — collector horticultural literature

Prior art & novelty

Is any of this new? A structured review of the peer-reviewed and germplasm literature (fig fingerprinting, the Son Mut Nou rimada work, sport/chimera mechanism, mislabeling rates) places the findings in context — and keeps the language honest about what would be a genuine first versus a confirmation.

      Under review — genotype vs. accepted variety status

      A grower review flags several of these clone calls as conflicting with accepted, phenotype-based variety status. Re-checked with stringent artifact filtering (high depth + allelic purity), the earlier ~3–5k "differences" collapse to essentially zero — true clones and the disputed pairs are indistinguishable from the same-tree control, while a genuinely distinct variety shows tens of thousands. So the conflict is real: the genome says clone; the phenotype says distinct.

      Evidence, not proof — and not a license to rename. A near-zero fixed-SNP distance shows a shared core genotype; from leaf DNA alone it cannot distinguish a genuine synonym from a chimeral/somatic sport. A physical resampling check found no sample mix-up, so these genome-identical calls are real — clonal synonymy and/or somatic sports, not lab errors. What leaf DNA still can't resolve is a change confined to the epidermis (L1), which contributes only a few per cent of leaf mass; the CDD Noir/Blanc color case, the Panachée / Blanca R / CDD Roja color trio and the Rimada placements are detailed in the lineage-puzzles write-up. ·

      Why the residual differences aren't real mutations

      ·

      ·

      Per-sample private differences

      Differences unique to one member of a trio — single digits each at the current floor, similar in the control and the clone group.

      Caveats & method

      Evidence, not proof. The same-tree control makes these clone calls robust, but formal cultivar-synonymy for a perennial crop should be confirmed by a qualified geneticist and by physical resampling. Leaf WGS captures the shared core genome, so a "0-difference" clone call and a chimeral sport can look identical here. Three of the mechanisms that could explain the striped-rimada and black/white-color phenotypes have since been tested directly rather than assumed away — structural variants genome-wide, hemizygous deletions via clone-mate ROH, and an L2-restricted (Pinot-gris-type) lesion via windowed depth and B-allele frequency — and all three are negative; see the cross-check section. What remains genuinely out of reach from leaf DNA is a change confined to the epidermis (L1), which is only a few per cent of leaf mass and so sits below the variant-detection floor, plus transposon insertions absent from the reference and any epigenetic difference. Those need tissue-specific sampling, long reads or methylation data. Stringent artifact filtering (high depth + allelic purity) reduced the disputed pairs to 0 clean differences (same as the same-tree control), and an independent GPU DeepVariant run has since reproduced the same pattern — so the clone calls are not a caller artifact.

      ·

      Alignment. NVIDIA Parabricks 4.7.0 fq2bam (GPU BWA-MEM + sort + mark-duplicates; no BQSR — there is no known-variant set for F. carica, so recalibration would treat real variants as errors) on a GB10, vs UNIPI Ficus carica v1.0 (GCA_009761775.1).

      Clonality metrics. Three, at increasing stringency: (1) whole-genome SNP distance over the alignment (clustering only, heterozygosity-inflated); (2) fixed (opposite-homozygote) differences over sites where both samples are confidently homozygous (the heatmap metric); (3) clean fixed differences after depth 15–150× + allelic-purity <5% filtering (the artifact-corrected biological floor). See "How to read the numbers".

      Tree/order. UPGMA on the fixed-difference matrix. The cohort tree is BIONJ/NJ on the per-comparable-site distance matrix; the branch-supported tree is the kinship suite's IQ-TREE GTR+ASC. FastTree is deliberately not used and its earlier agreement is not cited: it treats the alignment's IUPAC heterozygote codes as unknown characters and silently dropped ~25% of the matrix, so it was estimating from homozygous sites only — and in a highly heterozygous outcrosser shared heterozygosity is the clone signal.

      Provenance cross-check. Findings were compared against the Collector Fig Cultivar Encyclopedia, a traditional horticultural reference. It is used to corroborate names, phenotypes and provenance — not as a genetic authority — and it explicitly avoids research-derived identity claims.

      Would you like to have more figs sequenced? Please donate!

      Every accession in this panel had to be sequenced, and sequencing is the binding constraint on what this study can answer. The sharpest design limitations above are counts of one — one control tree, one cross-archive replicate, one first-degree reference pair — and each of those is relaxed the same way: another library. Every genome added also strengthens every call already made, because each new accession is another chance to falsify one.

      Support this work on Ko-fi

      ko-fi.com/persistentfig