Ruggedness
Is the landscape smooth or rugged?
A rugged landscape has many local peaks and abrupt changes in fitness between neighbors. Ruggedness makes it more likely that stepwise improvement stops short of the best candidate.
Explore ruggednessA fitness landscape maps the candidates in a design space, such as protein variants or reaction conditions, to their measured performance. GraphFLA builds this landscape from your data and characterizes its topography, revealing what makes a problem easy or hard to model and to optimize. These insights help you develop better predictive models and search algorithms.
$ pip install graphfla
20+
measures of landscape structure
7
landscape types for sequences, categorical, ordinal and mixed data
8
model landscapes and benchmark problems
9
step-by-step tutorials on published data
curated datasets from biology, chemistry and materials science
How it works
Illustrative data and profiles
Change amino acids at three sites to improve enzyme activity.
1Provide your data
A table of tested candidates, with one column for each variable and one for the measured property.
Your data
| Protein sequence | Activity |
|---|---|
| MKTAYIAKQRQISFVKSHFS | 23% |
| MKTAFIAKQRQISFVKSHFS | 41% |
| MKTAYIGKQRQISFVKSHFS | 36% |
| MKTAYIAKQKQISFVKSHFS | 57% |
| MKTAFIGKQRQISFVKSHFS | 42% |
| MKTAFIAKQKQISFVKSHFS | 64% |
| MKTAYIGKQKQISFVKSHFS | 73% |
| MKTAFIGKQKQISFVKSHFS | 91% |
2Build the landscape
GraphFLA links candidates that are adjacent in the design space, e.g., protein variants that differ by one amino-acid substitution, or reaction conditions that differ by one solvent.
GraphFLA
3Analyze its topography
Quantify peaks, ruggedness, neutrality and variable interactions.
Landscape profile
Local optima4 / 32
Roughness0.68
Neutrality19%
Fitness-distance correlation−0.25
Variable interactions
Choose a catalyst, solvent and reaction temperature to increase product yield.
1Provide your data
A table of tested candidates, with one column for each variable and one for the measured property.
Your data
| Catalyst | Solvent | Temp. | Yield |
|---|---|---|---|
| Nickel | Water | 40 °C | 32% |
| Nickel | Ethanol | 60 °C | 48% |
| Copper | Water | 60 °C | 54% |
| Copper | Ethanol | 80 °C | 81% |
| Palladium | Water | 40 °C | 65% |
| Palladium | Ethanol | 60 °C | 72% |
| Nickel | Acetone | 80 °C | 51% |
| Copper | Acetone | 40 °C | 74% |
2Build the landscape
GraphFLA links candidates that are adjacent in the design space, e.g., protein variants that differ by one amino-acid substitution, or reaction conditions that differ by one solvent.
GraphFLA
3Analyze its topography
Quantify peaks, ruggedness, neutrality and variable interactions.
Landscape profile
Local optima3 / 36
Autocorrelation0.61
Neutrality12%
Fitness-distance correlation−0.70
Variable interactions
Vary copper, zinc and nickel fractions to improve alloy hardness.
1Provide your data
A table of tested candidates, with one column for each variable and one for the measured property.
Your data
| Copper | Zinc | Nickel | Hardness |
|---|---|---|---|
| 60% | 25% | 15% | 145 HV |
| 60% | 20% | 20% | 170 HV |
| 65% | 20% | 15% | 158 HV |
| 65% | 15% | 20% | 185 HV |
| 70% | 20% | 10% | 135 HV |
| 70% | 15% | 15% | 162 HV |
| 55% | 25% | 20% | 176 HV |
| 55% | 20% | 25% | 198 HV |
2Build the landscape
GraphFLA links candidates that are adjacent in the design space, e.g., protein variants that differ by one amino-acid substitution, or reaction conditions that differ by one solvent.
GraphFLA
3Analyze its topography
Quantify peaks, ruggedness, neutrality and variable interactions.
Landscape profile
Local optima2 / 28
Autocorrelation0.83
Neutrality24%
Fitness-distance correlation−0.46
Variable interactions
Toggle caching, parallel builds and debug symbols to reduce build time.
1Provide your data
A table of tested candidates, with one column for each variable and one for the measured property.
Your data
| Cache | Parallel | Debug | Time |
|---|---|---|---|
| Off | Off | On | 92 s |
| On | Off | On | 61 s |
| Off | On | On | 55 s |
| Off | Off | Off | 78 s |
| On | On | On | 32 s |
| On | Off | Off | 49 s |
| Off | On | Off | 43 s |
| On | On | Off | 24 s |
2Build the landscape
GraphFLA links candidates that are adjacent in the design space, e.g., protein variants that differ by one amino-acid substitution, or reaction conditions that differ by one solvent.
GraphFLA
3Analyze its topography
Quantify peaks, ruggedness, neutrality and variable interactions.
Landscape profile
Local optima3 / 48
Autocorrelation0.72
Neutrality16%
Fitness-distance correlation−0.62
Variable interactions
Tune tree depth, forest size and minimum leaf size to reduce prediction error.
1Provide your data
A table of tested candidates, with one column for each variable and one for the measured property.
Your data
| Depth | Trees | Leaf | Error |
|---|---|---|---|
| 4 | 50 | 1 | 25.4 |
| 4 | 100 | 4 | 23.8 |
| 8 | 50 | 1 | 18.2 |
| 8 | 100 | 2 | 15.6 |
| 12 | 50 | 2 | 17.4 |
| 12 | 100 | 4 | 16.1 |
| 16 | 50 | 1 | 16.7 |
| 16 | 100 | 4 | 16.2 |
2Build the landscape
GraphFLA links candidates that are adjacent in the design space, e.g., protein variants that differ by one amino-acid substitution, or reaction conditions that differ by one solvent.
GraphFLA
3Analyze its topography
Quantify peaks, ruggedness, neutrality and variable interactions.
Landscape profile
Local optima2 / 36
Autocorrelation0.56
Neutrality21%
Fitness-distance correlation−0.39
Variable interactions
Quick start
Simply load your dataset, specify the variables and the objective, and build the landscape. The same steps apply to laboratory measurements, simulations and model evaluations.
$ pip install graphfla
import pandas as pdfrom graphfla import analysisfrom graphfla.landscape import ProteinLandscape df = pd.read_csv("variants.csv")X = df["sequence"] # Amino-acid sequence of each variantf = df["activity"] # Measured activity, to maximize landscape = ProteinLandscape(maximize=True)landscape.build_from_data(X, f)analysis.profile(landscape, seed=42)
import pandas as pdfrom graphfla import analysisfrom graphfla.landscape import Landscape df = pd.read_csv("reactions.csv")X = df[["catalyst", "solvent"]] # Reaction conditionsf = df["yield"] # Product yield, to maximize landscape = Landscape(maximize=True)landscape.build_from_data( X, f, data_types={c: "categorical" for c in X})analysis.profile(landscape, seed=42)
import pandas as pdfrom graphfla import analysisfrom graphfla.landscape import OrdinalLandscape df = pd.read_csv("alloys.csv") # columns: W, Re, Os, hardnessX = df[["W", "Re"]] # Os = 100 - W - Re, so it is impliedf = df["hardness"] # Hardness at 1,000 °C, to maximize landscape = OrdinalLandscape(maximize=True)landscape.build_from_data(X, f)analysis.profile(landscape, seed=42)
import pandas as pdfrom graphfla import analysisfrom graphfla.landscape import BooleanLandscape df = pd.read_csv("builds.csv")flags = ["gvn", "inline", "instcombine", "jump_threading", "sccp", "simplifycfg"]X = df[flags] # One on/off column per compiler flagf = df["time"] # Compilation time, to minimize landscape = BooleanLandscape(maximize=False)landscape.build_from_data(X, f)analysis.profile(landscape, seed=42)
import pandas as pdfrom graphfla import analysisfrom graphfla.landscape import Landscape df = pd.read_csv("grid_search.csv")X = df[["depth", "leaf", "features"]] # Hyperparametersf = df["rmse"] # Validation error, to minimize landscape = Landscape(maximize=False)landscape.build_from_data( X, f, data_types={c: "ordinal" for c in X})analysis.profile(landscape, seed=42)
Choose the class that matches your data.
What you can measure
Each feature is implemented from established studies in evolutionary biology and optimization. Together they let you explore your problem from complementary angles and understand it fully.
Is the landscape smooth or rugged?
A rugged landscape has many local peaks and abrupt changes in fitness between neighbors. Ruggedness makes it more likely that stepwise improvement stops short of the best candidate.
Explore ruggednessDo the variables act independently?
Variables interact when the effect of changing one depends on the values of others, known as epistasis in biology. Interactions that reverse the direction of an effect can block improving paths and create multiple peaks.
Explore interactionsCan local search find the global optimum?
Navigability describes how easily a search that accepts only improvements reaches the global optimum: from how many starting points, in how many steps, and whether fitness rises steadily towards it.
Explore navigabilityHow much of the landscape is flat?
Neutral changes leave fitness essentially unchanged. Large flat regions can stall a search, but they also connect distant candidates of equal quality.
Explore neutralityWhy landscape topography matters
Landscape features help explain why a predictive model or a search method works well on some problems but not on others. In the protein datasets below, compare each feature with the accuracy of zero-shot predictors and with the results of simulated directed evolution.
ProteinGym · Prediction performance across protein assays
Up to 480 evaluations · 10 runs per landscape
Performance
Every stage of landscape construction has been optimized in detail. The comparison below measures construction time and peak memory against a naive implementation that compares all pairs of candidates.
59,161×
faster construction at 1,048,576 candidates
2,384×
lower peak memory on the same build
Complete binary landscapes with 8 to 20 variables. Apple M4 Pro, Python 3.9, single thread. Peak memory includes Python and its dependencies. The baseline is a nested Python loop with a dense distance matrix.
Case studies
GraphFLA makes no assumption about what the variables or the objective represent. The same analysis applies to amino acids and reaction conditions, alloy compositions and neural architectures, drug doses and compiler flags.
Suzuki–Miyaura coupling
Ligand, base and solvent for one pair of reactants, scored by the UV signal of the coupling product.
384 reaction conditions
Electrochemical flow hydrogenation
Concentration, temperature and residence time in a flow reactor, scored by selectivity and current efficiency.
54 operating conditions
Tungsten–rhenium–osmium alloys
Tungsten, rhenium and osmium fractions of an alloy, scored by hardness at 1,000 °C.
496 alloy compositions
Hybrid perovskites · ABX₃
Organic ion, metal and halide of a hybrid perovskite crystal, scored by its calculated electronic band gap.
192 crystal compositions
BacPUS · Gut bacteria and polysaccharides
Pairs of gut bacterial strains and polysaccharide nutrient sources, scored by growth after 48 hours.
560 strain–substrate pairs
Cyanimide library · Mouse USP18
Amine and carboxylic-acid building blocks combined into inhibitors, scored by inhibition of mouse USP18.
7,504 building-block pairs
NCI-ALMANAC · Cancer cell assays
A partner drug for a fixed anticancer drug and the doses of both, scored by their effect on cancer-cell growth.
900 drug–dose combinations
NAS-Bench-201 · CIFAR-10
The operation on each of six connections in a neural-network cell, scored by image-classification accuracy.
15,625 network architectures
LLVM · Compiler flags
Ten LLVM compiler options switched on or off, scored by compilation time for a fixed workload.
1,024 compiler configurations
NeurIPS 2025 · Spotlight
ISSTA 2025
KDD 2025
IJCAI 2023