Navigation is translated. The technical pages are in English.
Computational biophysics and engine specifications

The discovery engine

A 5-stage computational pipeline. Public metagenomes are streamed into a local neural sieve, folded, and then simulated in explicit solvent — so candidates are ranked before any wet-lab spend.

2.3 TB
Metagenomic archive
Hundreds of millions of sequences evaluated
14,200/s
Sequences screened per second
41-head prefilter peak, 128-aa bucket
<15%
Sequence identity threshold
FoldTree 3Di remote homology
216K
Atoms in the largest solvated box
OpenMM TIP3P / TIP4P-Ice gate
01. Metagenome streaming 02. Neural sieve 02B. Biophysical calipers & triage 03. FoldTree & multi-modal console 04. Explicit OpenMM MD 05. CDMO/CRO transfer 06. Natural extremophiles
SECTION 01 Data acquisition and streaming

30TB of metagenomic sequence, streamed

Commercial enzyme discovery has historically drawn on a handful of lab-adapted production organisms in curated reference banks—a narrow slice of available genetic diversity. DarkBiome instead taps the full breadth of the public record: uncultured environmental metagenomes and sequenced isolate genomes, heavily weighted toward extreme physical habitats.

Habitats we draw from

Polar ice, snow and cold ocean water

Cold-adapted hydrolases and ice-associating proteins naturally functional at near-freezing temperatures (0–10 °C).

Hydrothermal and deep-sea sediments

Thermostable and barotolerant biocatalysts evolved to operate under high temperatures and extreme mechanical pressures.

Hypersaline and alkaline waters

Halotolerant and alkaliphilic enzymes that remain active and soluble in high-salt industrial brines without precipitation.

Contaminated and industrial sites

Specialized biocatalysts evolved under prolonged selective pressure to degrade synthetic chemical contaminants and xenobiotics.

The discovery corpus aggregates public sequence across uncultured metagenomic reads, assembled environmental contigs, isolate genomes, and specialized catalogs drawn from EMBL-EBI MGnify, NCBI, JGI, and UniProt. Our streaming engine indexes and triages sequences in memory, retaining qualified hits in a unified candidate registry of over 16.6 million records. To date, nearly 400 million distinct sequences have been evaluated, yielding nearly 1 million folded structural models.

SECTION 02 Neural architecture and local compute

Two neural tiers before folding

Screening and structure prediction have fundamentally distinct compute profiles, and run on optimized hardware tiers. The high-throughput neural screening tiers evaluate hundreds of millions of sequences on local Apple Silicon unified memory at negligible energy cost. Candidates that pass these filters advance to CUDA-accelerated batch folding and explicit-solvent molecular dynamics on GPU clusters (Section 04).

Ultra-high-throughput edge screening

By running optimized protein language representations directly on local edge neural silicon, DarkBiome screens incoming metagenomic sequences at tens of thousands of candidates per second with minimal energy draw.

  • • Tier 1 Multi-Target Sieve: Simultaneously scores sequences across all 35 active discovery profiles at rates exceeding 14,000 sequences per second.
  • • Tier 1.5 Target Disambiguation: High-resolution neural classifiers reject decoy homologs—separating ice nucleators from antifreeze proteins, and target esterases from unrelated CAZymes.
  • • 99.7% Sieve Efficiency: Eliminates non-viable sequences before committing GPU cluster hours to 3D structure prediction and all-atom simulation.
High-Throughput Funnel Architecture
Metagenomic Input Stream 397M+ Sequences
↓ Tier 1 Multi-Target Sieve (>14,000 seq/s)
Candidate Target Sieve 150M+ Evaluated
↓ Tier 1.5 Decoy & Homolog Rejection (<0.3% pass)
Qualified for 3D Folding Top Candidates

Clearing both neural tiers qualifies candidates for atomic-resolution structure prediction and biophysical diagnostics. By restricting 3D folding and all-atom molecular dynamics strictly to top-decile candidates, our pipeline delivers massive metagenomic exploration with exceptional capital and computational efficiency.

Biophysical gate 02A •

Three-tier secretion evidence ladder

Signal-peptide predictors have a well-known vulnerability: they frequently misclassify hydrophobic N-terminal transmembrane anchors as cleavable secretion signals. Expressed in an industrial host, such constructs remain membrane-bound and insoluble. We enforce rigorous topological verification:

Tier 1: Deep topology screening

Independent deep learning topology models (DeepTMHMM) identify authentic Sec/SPI cleavable signals while filtering out transmembrane spans and GPI anchors.

Tier 2: Cleavage boundary verification

Physicochemical boundary analysis enforcing strict hydrophobic core length and charge constraints to eliminate false-positive signal predictions.

Tier 3: Natural sequence integrity (ΔL = 0)

Zero chimeric end-caps or artificial mutations. Candidates are prioritized in their native evolutionary sequences for predictable expression and folding.

SECTION 02B Accelerated biophysical diagnostics & geometric triage

Sub-angstrom catalytic calipers and geometric triage

High computational folding confidence ensures a predicted backbone is coherent, but cannot guarantee catalytic readiness. Displaced active-site residues or collapsed coordination spheres silent-kill expression campaigns. Before advancing to molecular dynamics or synthesis, we subject candidates to sub-angstrom geometric validation across four key biophysical criteria.

Four biophysical diagnostic criteria

High-throughput geometric verification evaluated on folded structures prior to molecular dynamics.

01. Sub-angstrom catalytic calipers
Active-Site Triad Geometry

Measures precise inter-atomic distances between nucleophile, catalytic base, and catalytic acid residues. Candidates with misaligned catalytic machinery are automatically eliminated before wet-lab synthesis.

≤ 3.20 Å (Optimal) 3.20–3.80 Å (Relaxable) > 3.80 Å (Vetoed)
02. Metal coordination spheres
Metalloenzyme Verification

Calculates 3D coordination geometry around active-site metal centers (zinc, iron, manganese) to verify authentic pocket assembly and functional cofactor binding.

Geometry: Tetrahedral, Trigonal Bipyramidal, or Distorted / Vetoed
03. Aromatic percolation networks
Electronic & Structural Stacking

Maps aromatic ring packing networks (Phe, Tyr, Trp, His) to verify structural integrity and continuous charge-transfer pathways across bio-electronic scaffolds.

Benchmark: ≥ 70% continuous axial percolation
04. Biomaterial charge patterning
Phase Separation Profiling

Profiles sequence charge distribution and residue composition to predict biomaterial condensation and phase separation without forcing artificial globular folds.

Conformation: Globule, extended coil, or polyampholyte
Validation principles •

Eliminating common computational discovery traps

Physical active sites over proxies

Rather than relying on unconstrained sequence similarity or abstract scores, all candidates are evaluated against physical 3D catalytic geometries and explicit substrate-binding coordinates.

Transmembrane disambiguation

Hydrophobic transmembrane anchors frequently mimic secretion signals in naive predictors. Every candidate claiming extracellular secretion undergoes multi-tier topology verification to guarantee genuine solubility.

Archetype-specific representation

Intrinsically disordered proteins and flexible biomaterials are evaluated through sequence-charge profiling rather than forced into artificial globular predictions that fail in the wet lab.

SECTION 03 Structural phylogenetics and sequence novelty

FoldTree 3Di structural phylogenetics

Conventional sequence alignment fails below 20%–25% identity—the structural "twilight zone" where primary sequence diverges while 3D catalytic architecture remains intact.

Structural alignment across remote sequence space

By translating amino acid sequences into structural state alphabets (Foldseek 3Di), we construct structural phylogenetic trees. Candidates cluster by 3D active-site conformation rather than sequence identity alone.

Sequence BLAST (<15% identity)
Query: MKFLVLLVALVAAARAGDP...
Target: MQTVTRNLGLAAVTVSSAA...
Score: E-value = 8.42 (NO SIGNIFICANT HIT)

Homology missed. The enzyme is filed as an uncharacterized hypothetical protein.

FoldTree 3Di alignment (same fold)
3Di Q: vvvvvvvvddvvaavvvvvd...
3Di T: vvvvvvvvddvvaavvvvvd...
Score: TM-Score = 0.887 | RMSD = 1.12 Å

The fold aligns closely while the sequence does not. That is a measure of sequence novelty; it is where a freedom-to-operate analysis starts, not its conclusion.

Candidates share <15% sequence identity with incumbent reference enzymes while preserving catalytic active-site geometry, establishing structural homology with remote sequence divergence. Freedom to operate is assessed per asset and jurisdiction under definitive licensing diligence.

Partner diligence •

Interactive discovery console & diligence environment

DarkBiome provides commercial partners with a dedicated discovery interface for transparent structural diligence and downstream synthesis planning:

Interactive 3D structural cards

Examine high-resolution 3D tertiary models with automated camera orientation focused directly on catalytic clefts and active-site residues.

Biophysical & confidence telemetry

Live telemetry covering residue-level confidence, catalytic triad distances, surface hydrophobicity profiles, and multimer stability metrics.

Instant expression cassettes

One-click export of mature sequences with verified cleavage boundaries and automated codon optimization profiles for E. coli and Pichia pastoris.

Reproducibility & Audit Trail: All structural alignments, catalytic coordinates, and biophysical scores are deterministically versioned and verified against empirical reference controls to ensure absolute auditability during partner diligence.
✓ Verified candidate telemetry • Audit-ready provenance
SECTION 04 Thermodynamics and the explicit-solvent MD gate

The all-atom explicit-solvent MD gate

Predicted structures are vacuum snapshots. In an industrial reactor, enzymes face thermal agitation, shear, and pH swings. Before recommending synthesis, we subject candidates to all-atom molecular dynamics in explicit water to verify conformational stability under operating conditions.

Simulation protocol (OpenMM 8.0 / AMBER14 / TIP4P-Ice)

Solvation box
Up to 216k atoms

Explicit water eliminates vacuum artifacts and maps real solvent infiltration.

Backbone RMSD
1.14 Å – 1.31 Å

Rigid core preservation with loop breathing tracked under simulated process shear.

Reaction geometry
Collinear alignment

Active-site substrate orientation and catalytic triad geometry preserved throughout simulation.

Every candidate must clear four biophysical gates before partner delivery:

  • Solvation stability: Verifies active-cleft hydration without backbone collapse or solvent denaturing.
  • Catalytic triad persistence: Preserves sub-angstrom distance tolerances (≤3.20 Å) across production trajectories.
  • Solubility profile: Confirms high predicted aqueous solubility with zero solvent-exposed hydrophobic patches.
  • Aggregation resistance: Rejects motifs prone to forming inclusion bodies in fermentation hosts.

By screening out unstable conformers in explicit solvent before ordering synthesis, we prioritize candidates with verified catalytic geometry and high predicted solubility for partner expression.

Bench validation •

DarkBiome Daybreak

DarkBiome Daybreak is our bench software for macOS and Windows: one integrated tool for plate assays across campaigns, from ice nucleation to plastic clearance to thermal stability. Plate layouts carry campaign, candidate ID, and lot numbers, keeping all measured data linked to specific candidates.

Droplet freezing
Ice-nucleation proteins. Thermal video of 96- and 384-well plates: freezing detection, per-treatment Kaplan–Meier fraction frozen, K(T), T50, site densities and publication-ready reports.
Available
Turbidimetric clearance
PETase, polyamide and polyurethane enzymes, LPMOs, biomineralization. Substrate clearance (or precipitation) over time, from plate-reader absorbance CSVs.
Preview
Microcalorimetry
Carbonic anhydrase for DAC, methane monooxygenases. Reaction heat from per-well temperature series.
Preview
Thermal shift
Lanmodulin, rare-sugar isomerases. Melting temperature (Tm) screens from fluorescence-versus-temperature CSVs (DSF or qPCR instruments).
Preview

Instrument integrations parse standard plate-reader formats directly into candidate dossiers, with automated certificates of analysis generated for downstream diligence.

Biophysical gate 04A •

Hydration stability gate and constrained surface optimization

Structure predictors work in vacuum without explicit solvent, so a backbone that appears stable in isolation can unfold or aggregate in aqueous solution. Every candidate must pass an explicit-solvent hydration gate:

Explicit-solvent hydration relaxation

Each candidate is solvated in explicit water boxes and relaxed through energy minimization. Candidates exhibiting excessive conformational drift or core solvent penetration are identified as vacuum artifacts and removed prior to synthesis.

Pass threshold: Cα RMSD ≤ 2.5 Å | No core solvent infiltration
Constrained surface optimization

Where surface optimization is required to improve partner process compatibility, stochastic Monte Carlo refinement mutates only solvent-accessible residues while locking the packed hydrophobic core. Strict hydropathy and isoelectric constraints ensure high solubility and prevent aggregation.

Core strictly conserved | High intrinsic solubility
SECTION 05 How evaluation gets funded

Customer-funded CDMO/CRO transfer

DarkBiome owns no bioreactors and no pipetting robots. We license IP and buy synthesis. Under confidential evaluation packages and discovery partnerships under bilateral Delaware CDA/MTA, synthesis is commissioned with a certified contract research organization (Twist Bioscience, GenScript, Charles River) to produce the enzyme to order, with evaluation packages credited against commercial licensing.

Toll synthesis and blinded material transfer under Delaware CDA/MTA

STAGE 1: EVALUATION DELIVERABLES
  • ✓ Purified enzyme, 50 mg to 250 mg: synthesized to order, 8–12 weeks from order to delivery.
  • ✓ Evaluation package credited: 100% credited against upfront transfer on definitive license execution.
  • ✓ Bilateral Delaware CDA/MTA: includes a no-reverse-engineering covenant.
  • ✓ Benchmarking on your feedstock: tested on the partner's own waste streams.
STAGE 2: COMMERCIAL CONVERSION
  • ✓ Upfront technology transfer: $500,000 – $750,000 cash.
  • ✓ Development milestones: $1.25M – $2.0M, tied to pilot scale.
  • ✓ Running royalty: 4.0% – 5.0% of net sales, field-exclusive.
  • ✓ Sublicensing: 20% of non-royalty sublicensing income (NRSI).

The arrangement lets a partner measure catalytic performance on their exact feedstock before committing to tech transfer, and it keeps our incentives tied to what the enzyme actually does at their site.

SECTION 06 The thesis behind the pipeline

Why we discover from natural extremophiles

Over 99% of microbial life cannot be cultured under standard laboratory conditions. For decades, industrial biotechnology was confined to the tiny fraction of domesticated microbes that grow on agar plates, leaving vast planetary biochemical solutions untapped.

A few production strains against the whole public record

Industrial biotechnology spent four decades teaching a handful of lab-adapted bacteria to run chemistry they never evolved for.

The usual production strains
  • × E. coli, S. cerevisiae, B. subtilis
  • × Evolved for mammalian guts and temperate soils
  • × Enzymes denature at 45°C or at high ionic strength
  • × Need years of protein engineering to hold up at industrial scale
Public sequence (where we look)
  • ✓ Environmental metagenomes, plus sequenced isolate genomes
  • ✓ Evolved in cold, saline, alkaline and contaminated habitats
  • ✓ Dense surface salt bridges, and tolerance of high ionic strength
  • ✓ Already operating near industrial conditions, so no mutagenesis is needed

Our premise is that 3.8 billion years of selection has already produced catalysts for most of the chemistry industry wants, and that the work is finding them rather than designing them. Metagenome streaming, local neural filters and explicit-solvent dynamics are how we search; the output is a screened, patentable set of sequences.

Dataroom and technical diligence

See the underlying data

Molecular dynamics trajectories, RMSF loop fluctuation plots, ProstT5 3Di alignments and codon-optimized plasmid maps are available to prospective licensees under mutual NDA.

Request dataroom access → Review licensing terms