# Revilico product changelog and release notes
Source: https://docs.revilico.bio/changelog
Weekly release notes for the Revilico drug discovery platform: new engines, workflow improvements, billing changes, and bug fixes.
### New features
* Revamped the **Enamine Library** with a unified search and a streamlined review-and-submit flow, so you can move from compound discovery to quote submission in a single, guided workflow.
* Introduced **multi-match Enamine results** with a matching options panel, letting you compare several candidate matches per query and pick the best fit before requesting a quote.
* Added an **archive** action in [Compound Review](/docs/revcompound-review), so you can retire finished compound sets without losing the record.
* Added **bulk download** in [Compound Review](/docs/revcompound-review), letting you export selected compounds in one action.
### Improvements
* Enhanced Enamine Library quote submissions to carry richer compound and matching context end-to-end, giving reviewers and CRO partners more information at the point of decision.
* Extended shared-quote access so viewers can now download the associated compound list and preserve matching-mode context when following a shared link.
* Polished the Enamine Library UI for clearer navigation and result presentation.
### Bug fixes
* Fixed a **3D pose rendering issue** in [Compound Review](/docs/revcompound-review) that could prevent poses from displaying correctly.
* Corrected **dark-mode styling** in the MD Analysis and protein–water visualization views for readability parity with light mode.
* Resolved a small issue affecting certain [RevScreen](/docs/revscreen) runs, restoring expected behavior for affected pipelines.
### New features
* Introduced **pipeline sharing** in [Compound Review](/docs/revcompound-review), letting you share pipeline runs and their compound sets with teammates so reviews stay collaborative without re-uploading data.
### Improvements
* Refined [Compound Review](/docs/revcompound-review) backend workflows for more consistent behavior when sharing, tagging, and aggregating large compound sets.
### Bug fixes
* Fixed a **query issue in ensemble docking** ([RevScreen](/docs/revscreen)) that could return incomplete result rows for certain runs.
* Resolved a Compound Review issue that could prevent shared compound data from opening reliably for organizations.
### New features
* Launched a **CRO Directory** in the admin panel, giving admins a centralized view of contract research organizations to accelerate quote and outreach workflows.
* Added a **pipeline importer with drag-and-drop** in [Compound Review](/docs/revcompound-review), letting you seed reviews directly from prior pipeline outputs without manual re-entry.
* Introduced a **Resend Invite** action so admins can re-send user invitations without recreating them.
* Added the capabilities to review and orchestrate across several CRO providers to procure quotes and get wet lab work done.
### Improvements
* Enhanced [Compound Review](/docs/revcompound-review) with campaign tagging and like/dislike feedback on compounds, plus additional review workflow refinements.
* Added a **right-hand-side panel** for *My Experiments* and *Admin Quote* tickets so context stays visible while you work.
* Refreshed the sidebar with a **minimize toggle for the Credit Usage widget** and restricted the billing link to admins for cleaner navigation.
* Enforced a **credit cap on overage usage** to protect accounts from runaway consumption.
* Streamlined the metered credit configuration by removing legacy pricing multipliers, producing more predictable billing behavior.
* Improved **Adaptyv Bio** integration UX: the platform now live-validates stored API keys and surfaces missing or invalid credentials with clearer 401/403 error messaging across all Adaptyv Bio pages.
* Improved batching behavior in Compound Review for large compound sets.
### Bug fixes
* Fixed the **⌘/Ctrl + \[** shortcut so it now correctly collapses the left sidebar in classic navigation mode.
* Resolved a **pose-loading issue** in ensemble docking that could prevent results from rendering after a run.
* Fixed an issue where a feature flag was not applied correctly during pipeline execution.
### Metered Credit System
* Rolled out the Metered Credit System, moving from preview to full usage-based credit accounting for a clearer, fairer billing experience.
* Introduced Credit Bundles — prepaid and postpaid bundles to match usage patterns.
* Added a cap on overage credits to protect against unexpected usage spikes.
* Streamlined credit bundle options based on early user feedback.
### Compound Review
* Added Compare & Campaign to Compound Review — compare compounds side by side, with campaign tracking and like/dislike feedback.
### Platform
* Refreshed the sidebar interface for easier navigation.
* Delivered performance improvements across the platform.
### Billing
* Introduced the Metered Credit System (Preview), a refreshed credit experience ahead of full usage-based credits.
* Added Subscription Plan Controls to enable or disable subscription plans.
### Project Hub & Navigation
* Launched Project Hub Planning & Timeline — an interactive project timeline you can drag, resize, zoom, and jump to today.
* Added Classic Navigation Mode, letting users choose between classic (always-open) and compact side navigation.
### Collaboration
* Enhanced Team Chat with expanded Slack integration.
### Modeling & Data
* Expanded RevQSAR & MD Descriptor modeling and analysis support, with an updated interface.
* Added RevData Safeguards — checks when dropping data files into forms.
* Released a second set of Demo Pipelines for RevFEP, RevMut-PMX, and RevQS.
### Improvements & Changes
* Updated pipeline navigation to match new product routes.
* Enhanced the RevMD-Bind protein-ligand pipeline.
* Enhanced report generation.
* Delivered overall performance and responsiveness improvements across pipelines.
### RevFEP — Cloud-Native Free Energy Perturbation
* Launched RevFEP, a cloud-native Free Energy Perturbation engine for binding affinity prediction powered by OpenFE 1.10, OpenMM, and AWS Batch GPU execution.
* Added support for four calculation types: RBFE (relative binding free energy), ABFE (absolute binding free energy), and complementary protocols for lead-optimization ranking.
* Implemented seven advanced protocol stages including non-equilibrium MD, REST2 enhanced sampling, adaptive lambda optimization, and MM-GBSA pre-filtering.
* Automated the full pipeline from raw protein–ligand complex PDB inputs through structure preprocessing, alchemical network construction, GPU simulation, and statistical analysis with uncertainty quantification.
### RevQSAR — Quantitative Structure–Activity Relationship Modeling
* Released RevQSAR for downstream analysis of SMILES + property matrices generated across the platform.
* Added multiple featurization methods: Morgan Fingerprints (ECFP), MACCS Keys, GraphConv embeddings, and ChemBERTa embeddings.
* Added clustering modalities: K-means, hierarchical, spectral, and Tanimoto-based clustering, with UMAP, t-SNE, and PCA projections for chemical-space visualization.
* Enables substructure and scaffold discovery to guide generative chemistry and lead expansion.
### Documentation
* Published full RevFEP and RevQSAR documentation pages with workflow guides and example pipelines.
### Expanded Analysis & Virtual Biology
* Introduced RevTS for transition-state analysis.
* Brought the full RevSim Virtual Cell experience online (interface and analysis).
* Delivered the full RevSingleCell analysis experience.
* Added RevScreen Protein-Ligand Docking.
* Released the full RevQS quantum-structure experience and full RevMut-PMX mutation free-energy analysis.
* Added Data Export capabilities across modules.
* Expanded SDF V2000 file support in the molecule viewer, and added an element-change mapping option for RevFEP.
* Improved docking algorithms and default settings.
* Retired the older RevGEx and RevCluster experiments to streamline the platform.
### Free Energy & Modeling
* Shipped RevFEP V2, a major upgrade to free-energy analysis.
* Added RevMD-Bind Residence Time, a new residence-time MD analysis.
* Released RevQSAR AutoQSAR modeling, with new MD Descriptor Analysis for descriptor-level insights within RevQSAR / RevAnalytics.
### Virtual Biology & Structure
* Introduced RevQS quantum-structure analysis (backend and interface).
* Added RevGRN for gene-regulatory-network analysis.
* Brought RevSim Virtual Cell online with initial virtual-cell analysis capability.
* Expanded RevSingleCell with new single-cell analysis features.
* Released new demo pipelines across these capabilities.
### Interface
* Launched Dark Mode V2, a refreshed dark theme with updated styling and smoother animations.
* Enhanced the molecule viewer with integrated file downloads and the ability to download key analysis data.
* Added on-screen guidance when selecting the CHARMM36 force field for MMPBSA.
### Docking
* Set the default RevDock docking box size to 30 for better out-of-the-box results.
### Improvements & Changes
* Added a friendly reminder when trying to reuse a trial plan.
* Continued RevFEP analysis enhancements.
### RevFEP — Free-Energy Perturbation Launch
* Launched RevFEP, powered by OpenFE, with input presets for quick setup.
* Introduced RevBench for benchmarking across the platform.
### Project Hub & Collaboration
* Launched Project Hub to organize pipelines into projects.
* Added Team Chat with Slack channel integration.
* Rolled out Real-Time Updates across the platform.
* Released new demo pipelines for RevFEP, RevMut-PMX, and RevQS.
### Molecular Dynamics & Docking
* Shipped RevMD-Bind MD Insights, a new MDInsights analysis for protein-ligand simulations.
* Added RevMut-PMX mutation free-energy enhancements.
* Enhanced RevDynamics post-processing and improved ligand detection.
* Added support for multi-model protein structures.
* Added on-screen guidance in Rigid Receptor and flexible docking.
### Platform
* Made Central Hub noticeably faster.
* Added credit-system support for RevConformer and RevADMET.
* Updated feature display categories for easier navigation.
* Delivered general Central Hub improvements.
### Documentation & Guided Workflows
* Set up Revilico Documentation and Solutions workflows ([https://docs.revilico.bio/docs](https://docs.revilico.bio/docs)).
* Updated documentation implementations and integrated guided workflows into the right-hand side panel.
* Added demo pipelines across all features to preview generated data for each run.
* Added guided workflows in the right-hand side panel to assist users in conducting computational analyses.
* Fixed pipeline sharing issues to ensure all users can access generated data.
### Revilico Interpreter & Guide Revamp
* Revamped Revilico Interpreter and Guide for faster screen interpretation and seamless linking to the full documentation knowledge base.
* Added new models (GPT-5.2, Claude Sonnet 4.6) for smarter and more accurate interpretation.
* Improved Revilico Guide with integrated knowledge base querying and a transparent UI/UX.
* Integrated dual functionality of the Revilico Guide and Interpreter into a unified transparent UI with accessibility from the right-hand side panel.
### Docking & Virtual Screening
* Added Blind Docking and expanded box configuration in the Virtual Screening engine (Flexible Docking).
* Flexible Docking now supports multiplexed screens, enabling multi-receptor inputs against larger chemical libraries with CNN rescoring and pose refinements.
### Molecular Dynamics (MD)
* Improved MD simulation PDB extraction to create ensemble docking pipelines across multiple trajectories.
* Fixed issues related to erroneous structures generated during trajectory extraction.
* Enabled sub-pipeline termination across all MD features for multi-complex inputs, allowing users to stop full experiments across all sub-pipelines.
### Models & Large-Scale Pipelines
* Enabled large input set processing in Geometric Minimization and Thermochemistry workflows for scalable data abstractions.
* Stabilized ADMET, Pharmacophore Analysis, and Retrosynthesis capabilities for improved reliability and performance.
### Molecular Dynamics (MD)
* Upgraded MD engine to Revilico’s latest R\&D version across all MD features.
* Added support for 4-point water models: TIP4P, OPC, and OPC3 for improved simulation accuracy.
* Expanded Ligand Membrane MD simulations with additional lipids and membrane systems for permeability analysis.
* Delivered performance optimizations resulting in faster MD simulation execution.
### Simulation Configuration
* Added UI components to configure solvents and solvent concentrations across MD workflows.
* Included ion concentration configuration to better resemble biological conditions.
### Analysis & Visualization
* Introduced RMSF plots and RMSF trajectory visualization tools with spatial fluctuation color mapping across protein surfaces.
* Enhanced Protein–Ligand MD analysis with detailed comparison tables including MM(PB/GB)SA, decomposition, and PCA insights for improved free energy estimation across trajectories.
### Models & Pipelines
* Updated Geometry Minimization and Thermochemistry Neural Network Potential (NNP) model weights.
* Added solvent configuration support to resemble physiological conditions.
* Synchronized Boltz Co-Folding and Virtual Screening Docking pipelines with Protein–Ligand MD input sequences.
* Standardized download file naming for improved analysis and seamless workflows.
### Retrosynthesis
* Upgraded Retrosynthesis Engine to the latest Revilico R\&D version with architectural and performance improvements.
### Reliability & Operations
* Implemented automated Jira ticket creation on pipeline failures for improved incident tracking and operational visibility.
* New issues are automatically logged and triaged, enabling faster resolution within 24 hours.
### Docking & Pose Prediction
* Replaced classical empirical docking model with physics-based search algorithms powered by deep learning CNNs for more accurate poses and binding affinity calculations in Flexible and Ensemble Docking.
* Added the ability to extract receptor structures directly from Protein-in-Water MD simulations within defined timeframes and custom intervals for structural analysis across MD timescales.
* Introduced advanced configuration options to extract individual receptors at specific trajectory frame intervals for Ensemble Docking workflows.
* Enhanced Ensemble Docking workflows to allow grid box configuration across all trajectory structures.
### Visualization & UI
* Improved docking visualization UI to better identify and analyze binding sites and filter docking data effectively.
### Free Energy Calculations (Beta)
* Introduced ABFE and RBFE engines (Beta) enabling advanced alchemical transformations to compute ΔG binding of complexes.
### Platform & Architecture
* Implemented multiple stability enhancements across the platform.
* Implemented the Pharmacophore engine using Pritam Kumar Panda’s designs.
* Designed a new transcriptomic architecture for large-scale batch processing, improving scalability, reliability, and pipeline modularity for heavy workloads.
### Ligand Modeling
* Implemented ligand conformer search to generate low-energy 3D conformations, improving docking pose coverage and allowing chemists to evaluate energetic penalties across ligand–protein complexes.
# AlphaFold and OpenFold
Source: https://docs.revilico.bio/docs/alphafold-openfold
Protein Structural Prediction with Alphafold & OpenFold
## Why Use This Engine?
In the documentation below, we will use Revilico’s AlphaFold Engine and Openfold engine to design protein structures with high confidence for downstream applications (i.e. Pocket Identification, Docking, Molecular Dynamics Simulations, etc.). The core foundation of Computational Chemistry begins after you identify your target and generate its structure. Protein folding algorithms are such a large breakthrough because it allows us to now utilize structure based drug discovery approaches, enabling us to take a more targeted engineering approach to a previous meticulous guess and check process.
## Background
Protein structure determines function, binding sites, and druggability. A protein’s 3D shape dictates what molecules it can bind, what reactions it catalyzes, and whether it can be targeted by drugs. Without structural information, drug discovery relies on trial and error rather than rational, structure based design. At its root protein structure required experimental determination using X-ray crystallography, a process that was both costly, time consuming, and had a fairly large failure rate. This is where Alphafold comes into place. Alphafold is a deep learning system that predicts a protein’s 3D structure from its amino acid sequence by learning evolutionary patterns and physical constraints from solved structures, predicting accurate structures at a fraction of the cost traditionally required to render the structure experimentally.
In this guide, we will learn how to run AlphaFold, understand the theory behind it, and gain the intuition needed to derive novel insights from this pipeline. We will cover two Protein Folding engines on our platform, AlphaFold and OpenFold (a reiteration of AlphaFold with greater configurability suitable for research groups that need fine grained control over the prediction process rather than for production use. Simply, we have Alphafold2, an open-sourced version of the core model, and Openfold which resembles Alphafold3, a model traditionally reserved for enterprise.
**Structure Generation Workflows**
In order to run the AlphaFold Engine, we will do the following (1) name our pipeline (e.g. Pipeline #1), (2) upload protein sequence as csv or manual input, (3) configure the following parameters: Num Relax, Template Mode, MSA Mode, and Pair Mode. We will then run the pipeline, and open results once the pipeline has completed running. You can find your results in the Central Hub on the Command Center. The following workflow occurs on the backend.
The first step is sequence validation. It takes the amino acid sequence input and standardizes the input (e.g. checking for valid amino acid codes, removing white space, and validating sequence length). This is then passed to MSA generation where it searches sequence databases, aligns homologous sequences (e.g. aligning all the similar protein sequences based on their residue number), and produces Multiple Sequence Alignment (MSA) showing conservation (e.g. amino acid at a particular residue do not change across all similar protein sequences) and co-evolution (e.g. if position 10 mutates with a positive change, position 85 might compensate by mutating with a negative charge).
What this means is that a deep MSA will mean that we have high confidence in the structure as it shows consistency across multiple structures, and vice versa a shallow MSA will have low confidence as it does not have as many structures that are similar to it. If enabled, following the MSA step we will do the template search step, where it will search the PDB database for structures with sequence similarly and extract distance constraints from these structures, noting that these templates act as a soft hint rather than a hard constraint (i.e. a suggestion for how it should be structured rather than a command fixing the structure).
From there, given the context of MSA and Template search, features are then extracted and fed into a neural network, where 5 ranked models are generated, each with per-residue confidence (pLDDT), and Inter-residue confidence (PAE). Additionally with the num relax parameter, it will take the top n output structures (i.e. 0, 1, 5), and use the AMBER99SB force field to fix geometric issues (i.e. remove atom overlap, correcting bond angels, and optimizing side chain positions).
## Interactive AlphaFold Viewer
Explore AlphaFold protein structure prediction results in an interactive 3D viewer. View predicted structures colored by confidence (pLDDT), compare ranked models, and analyze per-residue quality metrics.
## Interactive OpenFold Viewer
Explore OpenFold protein structure prediction results in an interactive 3D viewer. OpenFold provides greater configurability for research groups needing fine-grained control over the prediction process.
# Boltz2
Source: https://docs.revilico.bio/docs/boltz2
Understanding Structure-Binding Dynamics using Revilico’s Boltz Co-Folding Engine
**Why Use this product?**\
Revilico’s Boltz Co-Folding engine is an AI powered tool that goes beyond structure prediction alone, also predicting binding affinity of protein-ligand or protein-protein complexes simultaneously. This engine enables rapid virtual screening, hit discovery, and ligand optimization for early-stage drug discovery without requiring expensive molecular dynamics simulations.
**Background**\
For most typical workflows, we would have to run our Alphafold Engine to generate our 3D protein structure, then use our pocket search engine to determine the pocket in which we want to bind to the ligand, then run our docking engine in order to find the binding affinity metrics as well as the best pose for the small molecules we would like to test.
This is a multistep process that can both take a significant amount of time and compute. This is where Revilico’s Boltz-Cofolding engine comes into place. This engine’s main goal is to provide fast, structure-level hypotheses for where a ligand is likely to bind, how the protein may fold or rearrange around the ligand, and assess whether the interaction looks plausible compared to other known protein-ligand complexes. If our goal is to quickly screen a protein and a set of ligands to see whether they should be sent into the wet lab, we can use Boltz Cofolding to confirm whether it has favorable conformations and activity.\
What is Boltz? Boltz is a machine learning based co-folding model trained on known protein ligand complex structures and functional activity scores. It will predict a single protein ligand complex geometry, confidence metrics for fold and interface, and provides the user a learned affinity signal.
Now we can dive into how this engine worksFirst we upload our Protein Sequences and our ligand SMILES strings. Our next step would be individual preprocessing of the protein sequences and the ligand SMILES strings, similar to how we operate the AlphaFold engine. Simply put, we will first validate the sequence (i.e. check that all amino acids are valid, assign residue indices and handle chain breaks if multiple proteins are present). We then undergo a Multiple Sequence Alignment (MSA) integration where MSA helps to align the sequence and structures to similar proteins, derived from evolutionary conservation of structure, allowing the model to see which residues have mutational and evolutionary conservation over time. To read more about how MSA works please refer to the Protein Folding Documentation here. Now for preprocessing the ligands, the SMILES strings are then converted to a graph structure where atoms are the nodes and bonds are the edges and each node is assigned features (i.e. element type, formal charge, hybridization, and aromaticity. These descriptors and graph representations of the ligands are then sent into an AI model to help predict the rest of the downstream outputs.
The next step would be merging the protein and ligand into a single system placing the protein residues and ligand atoms in the same graph representation, where there is no structure, only identities and relationships of the different input parameters. The job of the model is to then figure out how these different features can be represented together on a shared latent space (i.e. a representation of the multi-variable system with a lower dimension representation of vectors). Boltz algorithm takes the approach of learning the conserved underlying patterns of physics and chemistry that exist within the data rather than solving mathematical equations for binding interactions and energies at each step. Simply put, this means that the model will capture physics dynamics such as steric repulsion, favorable electrostatics, hydrophobic burial, and entropic preferences based on the experimental data that the model was trained on rather than having to calculate these dynamics at each step.
At this point within the algorithm, we have our network arrays of protein residues and ligand atoms representations/features. We will undergo a process of geometric optimization and iteration until we have reached our final optimized conformational state. The nodes in the graph each have their own properties (i.e. charge, hybridization, etc). The model will look at neighbors of atoms (nodes as represented by the algorithm), and evaluate how likely these nodes should be near each other within a specified conformation in a 3D predicted structure. Based on this likelihood, the nodes will update their internal representation accordingly and reevaluate. As the algorithm progresses and begins maximization of physical feasibility of certain conformations, interactions will increase their likelihood of being physically represented properly until the best conformation is submitted.
Lastly it is important to note that Boltz is an SE(3)-equivariant which means that when we rotate or translate the input, the output rotates or translates in the same way (i.e. the model does not align with absolute protein structure orientations, only relative geometry matters within the system. This is critical for physical realism and avoids any need for data augmentation as it guarantees that bond lengths stay consistent, angles behave properly, and structures obey spatial laws, all derived through MSA templates as a starting point.
## Interactive Results Viewer
Explore Boltz Co-Folding results interactively. View predicted protein-ligand complex structures with pLDDT confidence coloring, pTM/ipTM metrics, and binding affinity predictions.
# BoltzGen
Source: https://docs.revilico.bio/docs/boltzgen
Understanding Joint Folding and Binding Dynamics using Revilico’s BoltzGen Co-Folding Engine
## Why Use this product?
The BoltzGen Cofolding Engine is a diffusion based tool that generates thousands of candidate binders with validated binding poses enabling rapid discovery of high affinity binders for drug development, diagnostics, or protein engineering applications without requiring experimental screening. This tool is best used when you need to design novel protein or peptide binders against therapeutic targets through AI powered de novo generation that simultaneously optimizes binding affinity, structural stability, and sequence diversity.

## Background
Typical workflows for discovering whether a molecule will bind to a protein pocket are a multistep process that is both time consuming and computationally expensive. And even with these workflows it is not guaranteed that the pocket we select is the correct pocket. BoltzGen Co-Folding is a tool that can be deployed to design a new binder that will both fold into a realistic 3D structure and bind to a target in a realistic 3D pose, in one unified generative process. This process can bypass the usual process of protein folding, then pocket search, then docking for traditional small molecule therapeutic developments.
BoltzGen Co-Folding works by uploading a target structure, typically a fixed protein structure as a PDB file, then design constraints for the binder in which we want to create (e.g. 80..140 is a length constraint for the number of residues). BoltzGen will then take these two structures and build one joint complex where we have the fixed target, and the designed binder (i.e. small molecule, peptide, etc.). BoltzGen is usually utilized for designing biologics modalities for therapeutics against specific targets.
BoltzGen will then run a generative model that is an all-atom diffusion style generative model, where designed entities are represented at the atomic level, grouped into tokens (i.e. residues/nucleotides), while designed parts are represented with naked residue/atom types in a fixed-size representation. The model will take in an input of known atoms/coordinates for the target, tokens for all residues (target + binder), with masks indicating which residue/atom types are designable, and conditioning signals based on the design constraints we have provided in the input. It will output a generated all-atom 3D structure for the whole complex, where a denoiser prediction will be used to iteratively refine the sample. In order words, the model is sampling a full complex from a learned distribution to see what real complexes look like, conditioned on your target and constraints. It can be denoted by the following equation:
$$
P(\text{sequence}, \text{structure}, \text{pose} \mid \text{fixed anchor}, \text{constraints})
$$
The model starts from a highly random complex and iteratively transforms it into a more statistically realistic protein-target complex, according to patterns learned from real structures. This concept of a diffusion style generative model takes on the idea of taking a protein complex adding noise until it looks like random junk, then training a neural network to undo this noise towards an optimized 3D conformation and engagement. The goal is to create a model that can reliably clean up noise resulting in a realistic structure. This can be denoted by the following equation:
$$
D_\theta(x_t, t, \text{conditioning})
$$
Where $x_t$ is the noisy version of a structure at step $t$, $t$ is how noisy it is, and conditioning is the fixed protein, ligand, constraints. Our output will be a less noisy structure.
At the generation step BoltzGen will use a probability flow ODE to generate the new complex. It can be denoted by the following equation:
$$
\frac{\partial x}{\partial t} = - \frac{x - \mu_\theta(x,t)}{t}
$$
Where X is the current molecular configuration at a given state, t is a continuous parameter controlling how noise the structure is with large t indicating very noisy random structure and small t indicating low noise and close to a final realistic structure, μ*θ(x*,t) is the model’s denoised prediction which represents the model’s best guess of what this structure would look like if it were clean and realistic, and dx/dt indicating the direction of change of how to update the structure as noise is reduced.
To summarize, this model will deploy design steps to iteratively transform a random joint structure into a more statistically realistic bound complex, where the design constraints and the presence of the fixed target shape what ‘realistic’ means at every step
## Interactive Results Viewer
Explore BoltzGen Co-Folding results interactively. View generated protein designs with filtering criteria, aggregate statistics, ranked structures, and metrics visualization.
# 3D Visualization
Source: https://docs.revilico.bio/docs/rev3d-viz
Interactive 3D Molecular and Structural Visualization with Mol* for Proteins, Ligands, Densities, and Trajectories
## Why Use 3D Visualization?
3D Visualization is Revilico's interactive structure viewer, built on Mol\*, for exploring protein structures, protein-ligand complexes, molecular dynamics trajectories, and electron density maps directly in the browser. It is the primary tool for visual inspection of computational results from RevTarget, RevBind, and RevDynamics engines, enabling researchers to examine binding poses, assess structural quality, and communicate structural findings without requiring a local molecular graphics installation.
## Background
Mol\* (pronounced "molstar") is a modern, high-performance molecular visualization framework developed by the RCSB PDB and PDBe teams. It renders large molecular assemblies efficiently in the browser using WebGL, supports a wide range of structure and data formats, and provides a comprehensive set of representation and analysis tools. RevTarget integrates Mol\* as the primary visualization layer for all structural outputs across the platform.
## Interface Panels
**Left panel (data and options):**
The Home tab provides structure loading tools. Structures can be loaded from PDB by entering a PDB ID with download source selection (PDB or PDB IDs), or downloaded as PDB files. Additional options include:
* **Download Density:** Fetch electron density maps for deposited structures
* **Download File:** Open local structure files
* **Open Files:** Access files from the Data Engineering environment
* **Load Trajectory:** Load MD trajectory files for dynamic visualization
* **Load Genome 3D (G3D):** Load genome 3D structure data
* **Download Tunnels:** Overlay channel and tunnel calculations
* **Zenodo Import:** Import structures and data from Zenodo repositories
* **Remote States:** Access saved visualization states
**Main canvas:**
The central 3D canvas renders the molecular structure with full rotation (drag), zoom (scroll), and translation (right-drag) controls. The canvas supports simultaneous display of multiple structure components, density maps as isosurface meshes, and trajectory playback.
**Right panel (Structure Tools):**
The Structure panel provides representation and analysis controls:
**Measurements:**
Add distance, angle, dihedral, and orientation measurements directly on the 3D structure by clicking atoms. Measurements persist and can be toggled or deleted.
**Quick Styles:**
Apply pre-configured representation presets with a single click:
* **Default:** Standard cartoon backbone with ball-and-stick ligand
* **Cartoon:** Cartoon secondary structure representation with colored helices, sheets, and loops
* **Spacefill:** Van der Waals sphere representation for each atom
* **Surface:** Molecular surface (solvent-accessible or solvent-excluded)
**Apply Style:**
* **Default:** Standard atom coloring by element and chain
* **Illustrative:** Cel-shading style for publication-quality renders with flat colors and visible outlines
**Apply Wiggle (Animation):**
Simulate thermal motion by applying positional displacement to atoms. Dynamics and uncertainty controls adjust the amplitude and visual feedback of the wiggle animation.
**Components:**
Fine-grained control over individual structural components. Each component (polymer chain, ligand, water, ions) can be shown or hidden, and its representation and color scheme can be independently configured using the add and delete controls.
**Export:**
* **Export Models:** Download the currently loaded structure as a PDB or mmCIF file
* **Export Animation:** Render and download a trajectory animation as a video file
* **Export Geometry:** Export the rendered 3D geometry for external visualization tools
## Running the Engine
### Inputs
| Input | Description |
| --------------- | ---------------------------------------------------------- |
| PDB ID | Fetch directly from RCSB or PDBe by 4-character identifier |
| Structure file | PDB, mmCIF, BCIF, SDF, MOL, XYZ, GRO formats |
| Trajectory file | DCD, XTC, TRR, PDBQT formats from MD simulations |
| Density map | CCP4, MAP, MRC, DSN6 electron density formats |
| Topology file | PSF, TOP, ITP, PRMTOP for trajectory context |
### Outputs
* **Interactive 3D view:** Rotatable, zoomable molecular visualization with all loaded components
* **Measurements:** Annotated distance, angle, and dihedral values on the structure
* **Exported structure:** PDB or mmCIF file of the current state
* **Exported animation:** Video file of trajectory playback
* **Exported geometry:** 3D mesh for external rendering
# RevADMET
Source: https://docs.revilico.bio/docs/revadmet
Assessing Compound Developmentality Risk with Revilico’s ADMET AI Engine
## Why Use this product?
ADMET-AI rapidly predicts 41 pharmacokinetic and toxicity properties, including absorption, distribution, metabolism, excretion, and toxicity endpoints, using machine learning models trained on Therapeutics Data Commons datasets. This enables drug discovery teams to filter thousands of compounds from virtual screening or generative AI in minutes, eliminating molecules with poor drug-like properties (low solubility, hERG liability, hepatotoxicity) before expensive synthesis and testing. By providing fast, accurate ADMET predictions with the highest performance on industry benchmarks, ADMET-AI bridges the gap between computational hit identification and experimental validation, dramatically reducing the cost and time of lead optimization for later stage tasks after activity optimizations.

## Background
Whether the ligand will engage with and bind to the protein is only one part of the question when it comes to whether a drug can be deemed effective and be passed into the clinic. Predicting the compound properties answers the other part of the question. It answers whether the drug will deliver the proper effect while maintaining other important properties and minimizing toxicity. By creating a high throughput screening system that can predict the properties of a specific compound, chemists will be able to prioritize molecules with a higher chance of succeeding in development.
Introducing Revilico’s ADMET AI engine, it is a data driven machine learning model, trained on large experimental ADMET datasets that is designed to rapidly assess developability risks of small molecules solely based on their chemical structure, enabling users to assess compounds based on absorption, metabolism, distribution, excretion, and/or toxicity before synthesis and experimental testing.
Now how does it work? This model is based on a principle that many ADMET outcomes are strongly correlated with high-level chemical structure and physicochemical properties rather than detailed molecular mechanisms (i.e. Membrane permeability correlates with size and polarity). Rather than calculating these properties using extensive experimentation, ADMET-AI deploys a machine learning model that can learn these properties from the data. First we begin with our user input, or SMILES strings. We then ask the question, how can we transform this string into something the model can understand and extract value from. We understand the principles that the behavior of an atom depends on its neighbor, functional groups behave differently depending on scaffold context, and small structural changes can lead to large biological effects. Using this principle of the relationships between atoms and the graph like nature of a molecule on a 2D/3D coordinate system, we can deploy a Graph Neural Network (GNN), where the model learns an atomistic representation of its local chemical environment including nearby functional groups, electronic context, and structural constraints imposed by the scaffold. Here we can produce a value or embedding in which the model can understand.
The model itself was trained on ADMET datasets that contain a wide range of small molecules with known outcomes measured from in vitro assays, in vivo studies, and biochemical screens. During training, the model learns a mapping from molecular embeddings to ADMET properties. At inference time, the input molecule’s embedding is passed through this learned mapping to generate predictions.
## Interactive Results Viewer
Explore ADMET AI prediction results interactively. Select a molecule to view its predicted properties across absorption, distribution, metabolism, excretion, and toxicity categories.
# RevAgent
Source: https://docs.revilico.bio/docs/revagent
Your AI Science Partner — An Intelligent Agent That Understands Your Science, Orchestrates Your Workflows, and Accelerates Discovery Through Natural Conversation
## Why Use RevAgent?
RevAgent is Revilico's AI research partner, purpose-built for drug discovery workflows. Rather than navigating individual engines manually, RevAgent understands your scientific question in natural language, selects the appropriate computational tools, executes them in sequence, and synthesizes the results into a coherent scientific narrative. It functions as a senior research collaborator that never forgets context, can orchestrate multi-step analyses across the entire Revilico platform, and explains its reasoning at every step.
## Background
Large language models with tool-use capabilities can serve as intelligent orchestration layers in scientific workflows, translating natural language research questions into concrete computational tasks and interpreting results in biological and chemical context. RevAgent is powered by Claude Sonnet and has native access to Revilico's full engine suite, enabling it to reason about molecular structures, genomic data, simulation results, and experimental readouts within a single conversation thread.
The interface is divided into two panels that together provide transparency into both the scientific reasoning and the execution mechanics.
## Co-pilot Panel
The Co-pilot panel on the left is the primary interaction surface. It displays RevAgent's scientific responses in natural language, synthesizing results from tool calls, interpreting data outputs, and providing mechanistic explanations grounded in the biological and chemical context of the query. Responses are structured to be directly actionable, highlighting key findings, flagging caveats, and suggesting next steps in the research workflow.
The input field at the bottom accepts free-text queries. Example queries the agent is designed to handle include: analyzing a drug-wellness profile, comparing binding affinity predictions across docking methods, interpreting differential gene expression results, or designing a lead optimization strategy for a specific chemical scaffold.
Files from the Data Engineering environment can be dragged directly into the input field, enabling the agent to reason over uploaded datasets, molecular files, or analysis outputs without requiring manual data transfer between tools.
**Model:** Claude Sonnet 4.5 (configurable via the model selector in the input bar).
## Executor Panel
The Executor panel on the right shows the step-by-step task execution log in real time. When RevAgent decides to invoke a computational tool, the Executor displays each step as it runs: which engine is being called, what parameters are being passed, what the intermediate output is, and whether the step succeeded or requires adjustment. This panel provides full transparency into the agent's reasoning chain, enabling researchers to audit the workflow, identify which step produced a given result, and intervene if a step needs to be modified.
The Executor is idle when no task is running, showing a "Waiting for task" state. Once a query is submitted, steps appear sequentially as the agent works through the analysis.
## Workflow
1. Enter a research question in the Co-pilot input field, optionally attaching files from Data Engineering
2. RevAgent parses the scientific intent and constructs an execution plan
3. The Executor panel shows each tool invocation in real time as the plan runs
4. The Co-pilot panel displays the synthesized scientific response with key findings, interpretations, and suggested next steps
5. Follow-up questions can be asked in the same conversation thread, with the agent retaining full context of all prior steps and results
## Running RevAgent
### Inputs
| Input | Description |
| ---------------------- | ---------------------------------------------------------------------------- |
| Natural language query | Research question, analysis request, or scientific hypothesis in plain text |
| Attached files | Molecular files, datasets, or analysis outputs dragged from Data Engineering |
| Model selection | Claude model version (configurable in the input bar) |
### Outputs
* **Co-pilot response:** Scientific interpretation, key findings, mechanistic context, and next-step recommendations
* **Executor log:** Step-by-step record of all tool invocations, parameters, and intermediate outputs
* **Generated artifacts:** Any plots, tables, or structured data produced by invoked engines are embedded in the response
# RevBench
Source: https://docs.revilico.bio/docs/revbench
Benchmarking Docking and Co-Folding Predictions Against Experimental Bioactivity Data
## Why Use This Engine?
In the documentation below, we will use Revilico's RevBench engine to evaluate how well computational docking and co-folding predictions align with experimental bioactivity measurements. RevBench automates the full benchmarking pipeline: retrieving experimental assay data from public databases, merging it with prediction outputs, and computing a comprehensive set of statistical metrics that reveal where a given computational method succeeds or fails for a target of interest.
## Background
Virtual screening methods such as docking and co-folding produce ranked lists of compounds with associated predicted scores. Evaluating whether these rankings meaningfully reflect experimental binding data is a non-trivial challenge. Experimental IC50, Ki, and Kd values are measured across heterogeneous assay conditions, stored in multiple public databases, and distributed at different scales. RevBench addresses this by standardizing the benchmarking workflow around a single protein target defined by its PDB ID, collecting bioactivity data from ChEMBL, PubChem, BindingDB, and IUPHAR, and computing statistically rigorous correlation, discrimination, and enrichment metrics against any prediction input the user supplies.
The platform supports two prediction modalities. Boltz2 co-folding predictions provide a predicted pIC50, a confidence score, an affinity probability, and an interface pTM (ipTM) score per compound. AutoDock Vina docking predictions provide binding free energy estimates (kcal/mol) from static, flexible, and ensemble docking runs. Both modalities are evaluated on the same experimental benchmark set, enabling direct and consistent comparison.
## Target Setup and Structure Preparation
The benchmarking workflow begins with a PDB ID. The engine retrieves target metadata from RCSB PDB, UniProt, ChEMBL, and NCBI, downloading the co-crystal structure and extracting the primary protein chain and the co-crystal inhibitor as separate PDB files. The docking box is computed automatically from the ligand centroid, extending by 10 Angstroms in each dimension to a cubic grid suitable for AutoDock Vina configuration.
## Experimental Dataset Assembly
Bioactivity data is retrieved from four databases: ChEMBL (paginated activity records by target ChEMBL ID), PubChem (assay results by gene ID), BindingDB (ligand affinities by UniProt ID, preferring isomeric SMILES), and IUPHAR/GtoPdb (ligand interactions, with pKd/pKi/pEC50 converted to nanomolar units via $\text{nM} = 10^{9 - p}$).
All records are deduplicated to one row per unique SMILES, prioritizing IC50 over Ki and Kd when multiple assay types exist for the same compound. Activity labels are assigned by the following rules in priority order: if the database provides an explicit active or inactive outcome it is used directly; the compound is labeled inactive if comments indicate no activity, if IC50 exceeds 100 micromolar, if HTS percent inhibition is below 10%, or if biophysical comments indicate no binding; otherwise the compound is labeled active. The final benchmark set uses a 9:1 inactive-to-active ratio for enrichment analysis to match industry-standard virtual screening conditions.
## Benchmarking Metrics
**Parity Plot and Correlation Analysis**
Experimental pK values are computed from quantitative measurements as:
$$
\text{pK} = -\log_{10}\left(\text{value}_{\text{nM}} \times 10^{-9}\right)
$$
Predicted scores are aligned to the same scale. For Boltz2, the reported pIC50 is used directly. For Vina static docking, binding free energy in kcal/mol is converted to an approximate pK unit using the linear relationship $\text{pK} = (\Delta G + 5.89) / 1.364$.
Five correlation metrics are reported with 95% bootstrap confidence intervals (1,000 resamples):
**Pearson r** measures linear correlation between predicted and observed potency:
$$
r = \frac{\sum (x_i - \bar{x})(y_i - \bar{y})}{\sqrt{\sum (x_i - \bar{x})^2 \sum (y_i - \bar{y})^2}}
$$
**Spearman rho** measures rank-based correlation and is robust to outliers.
**Kendall tau** measures rank concordance across all compound pairs.
**RMSE** and **MAE** quantify prediction error magnitude:
$$
\text{RMSE} = \sqrt{\frac{1}{n}\sum_{i=1}^{n}(y_i - \hat{y}_i)^2}, \quad \text{MAE} = \frac{1}{n}\sum_{i=1}^{n}|y_i - \hat{y}_i|
$$
**Lin's Concordance Correlation Coefficient (CCC)** combines precision and accuracy, penalizing predictions that correlate well but are systematically shifted:
$$
\text{CCC} = \frac{2\rho \sigma_x \sigma_y}{\sigma_x^2 + \sigma_y^2 + (\mu_x - \mu_y)^2}
$$
**Enrichment Factor**
The enrichment factor (EF) at a given fraction $\chi$ measures how efficiently a virtual screening ranking recovers known actives compared to random selection:
$$
\text{EF}(\chi) = \frac{\text{fraction of actives in top } \chi\text{\%}}{\chi}
$$
An EF of 1.0 corresponds to random retrieval. EF values at 1%, 5%, and 10% of the library are reported with bootstrap uncertainty bands.
**ROC Curve and AUC**
The ROC curve plots true positive rate (TPR) against false positive rate (FPR) as the score threshold is varied. For Boltz2, two ROC curves are generated separately using affinity probability and predicted pIC50 as the ranking signal. AUC of 0.5 indicates random discrimination; AUC of 1.0 indicates perfect separation of actives from inactives.
**Confidence Calibration**
For Boltz2 predictions, confidence score and ipTM are assessed against per-bin prediction accuracy. Compounds are binned by confidence signal value (0 to 1 in 0.2-width bins), and within each bin the fraction of predictions with absolute error below 0.5 pIC50 units (the success rate) and the mean absolute error are reported. Well-calibrated models show monotonically increasing success rates with increasing confidence.
**Correlation Heatmap**
A pairwise correlation heatmap is computed across all numeric columns from the merged prediction-experiment dataset. Pearson r, Spearman rho, Kendall tau, R-squared, and CCC are all available as the color metric. Cells with fewer than five paired observations are left blank.
## ADMET and Structural Alert Profiling
In addition to benchmarking predictive accuracy, RevBench profiles the benchmark compound set for drug-likeness and structural liabilities using RDKit. Computed descriptors include molecular weight, LogP, hydrogen bond donors and acceptors, topological polar surface area, rotatable bonds, ring counts, QED (quantitative estimate of drug-likeness), and estimated water solubility (ESOL model). Drug-likeness rules assessed include Lipinski Rule of Five, Veber criteria (rotatable bonds and TPSA), and Ghose filter. Structural alerts are flagged against the PAINS, Brenk, and NIH MLSMR catalogs to identify compounds prone to assay interference or containing known problematic substructures.
## Running the Engine
### Inputs
| Input | Required | Description |
| -------------------- | -------- | ----------------------------------------------------------------------------------------------- |
| PDB ID | Yes | 4-character PDB identifier of the target |
| Boltz2 output | Optional | CSV with `ligand_smiles`, `predicted_pic50`, `confidence_score`, `affinity_probability`, `iptm` |
| Vina static output | Optional | CSV with `ligand_smiles`, `best_affinity`, `mean_affinity` (kcal/mol) |
| Vina flexible output | Optional | CSV with per-pose ΔG and CNN affinity/pose scores |
| Vina ensemble output | Optional | CSV with per-conformation ΔG values across receptor snapshots |
### Outputs
* **Experimental benchmark set:** Deduplicated active/inactive compound set with source database, SMILES, assay type, and measured value
* **Parity plots:** Predicted vs. observed pK scatter plots with OLS regression line and 1-sigma confidence band
* **Correlation metrics table:** Pearson r, Spearman rho, Kendall tau, R-squared, CCC, RMSE, MAE, and bias with 95% CI
* **Enrichment curves:** EF at 1%, 5%, 10% with bootstrap uncertainty bands
* **ROC curves:** With AUC and 95% CI
* **Calibration plots:** Confidence bin vs. success rate and MAE (Boltz2 only)
* **Correlation heatmap:** Pairwise correlations across all numeric prediction and experimental columns
* **ADMET table:** Per-compound physicochemical descriptors, drug-likeness rule results, and structural alerts
* **Docking-specific plots:** CNN vs. physics affinity scatter, per-compound pose distribution boxplots (ensemble), active/inactive score distribution histograms
# Cell Death
Source: https://docs.revilico.bio/docs/revcell-death
Apoptosis and Necrosis Prediction Across Cancer Cell Lines from Drug-Induced Stress
## Why Use This Engine?
In the documentation below, we will use Revilico's Cell Death assay engine to predict the mode and magnitude of drug-induced cell death across a cancer cell line panel. Rather than measuring bulk viability, this assay decomposes the death response into apoptotic and necrotic fractions and reports caspase 3/7 activity, providing mechanistic insight into how a compound kills cells and whether it engages the programmed apoptotic pathway.
## Background
Cell death from cytotoxic drugs occurs through two primary mechanisms. Apoptosis is a programmed, energy-dependent process characterized by caspase activation, membrane blebbing, and ordered DNA fragmentation. It is the preferred mechanism for cancer drugs because it is immunologically quiet and does not trigger inflammation. Necrosis is a passive, uncontrolled death caused by membrane rupture and organelle swelling, typically associated with high-dose toxicity or cellular stress exceeding the capacity for organized apoptotic signaling. Distinguishing these modes is clinically relevant: drugs that induce predominantly apoptotic death are generally better tolerated than those that induce necrosis.
RevAssay's Cell Death module wraps the Revilico Virtual Cell GNN+MLP viability model (R-squared 0.87, trained on GDSC1+2 and DepMap CCLE 24Q4) to derive a cellular stress signal, then applies mechanistic scaling to produce apoptosis percentages, necrosis percentages, and caspase 3/7 activity readouts that reflect the dose-dependent induction of each death pathway.
## Simulation Model
The viability prediction from the GNN+MLP model is converted to a cellular stress signal:
$$
\text{stress} = 1 - v(c)
$$
Where $v(c)$ is the predicted fractional viability at concentration $c$. The apoptotic and necrotic fractions are derived from this stress signal with pathway-specific scaling factors:
$$
\text{Apoptotic\%} = \text{stress} \times 74 + \varepsilon_a, \quad \varepsilon_a \sim \mathcal{N}(0, 6^2), \quad \text{clamped to } [0, 95]
$$
$$
\text{Necrotic\%} = \text{stress} \times 14 + \varepsilon_n, \quad \varepsilon_n \sim \mathcal{N}(0, 4^2), \quad \text{clamped to } [0, 25]
$$
$$
\text{Viable\%} = 100 - \text{Apoptotic\%} - \text{Necrotic\%}
$$
The scaling coefficients reflect the biology of EGFR-driven and generic cytotoxic stress: the majority of stress-induced death in epithelial cancer cell lines occurs through the intrinsic apoptotic pathway (caspase 9 then caspase 3/7 activation), with necrosis emerging only at extreme stress levels. Caspase 3/7 bioluminescence is modeled as proportional to the apoptotic fraction:
$$
\text{Caspase RLU} = \text{Apoptotic\%} \times 11 + \varepsilon_c + 5, \quad \varepsilon_c \sim \mathcal{N}(0, 15^2)
$$
The intercept of 5 RLU represents baseline caspase activity in untreated cells.
## Parameters
| Parameter | Default | Description |
| ---------------------- | -------- | --------------------------------------------------- |
| Death mode | Combined | `Apoptosis`, `Necrosis`, or `Combined` display mode |
| Caspase detection | On | Enable or disable caspase 3/7 RLU readout |
| Exposure hours | 72 h | Duration of drug treatment (matches GDSC baseline) |
| Hill coefficient | 1.2 | Dose-response curve steepness |
| Biological variability | None | Noise level applied to replicate predictions |
| Replicates | 3 | Number of replicate wells for CI computation |
## Outputs
* **Apoptotic %:** Fraction of cells undergoing programmed apoptotic death at each concentration
* **Necrotic %:** Fraction of cells undergoing uncontrolled necrotic death
* **Viable %:** Fraction of surviving cells
* **Caspase 3/7 RLU:** Luminescence proxy for executioner caspase activation
* **Dose-response curves:** Per-cell-line curves for each readout across the concentration range
* **96-well heatmap:** Plate-view of apoptotic or viable fraction colored by death mode
* **Stacked bar charts:** Apoptotic (orange), necrotic (red), and viable (green) fractions per cell line per concentration
# Compound Design
Source: https://docs.revilico.bio/docs/revcompound-design
Interactive 2D Molecular Structure Drawing and Editing with Direct Export to Revilico Engines
## Why Use Compound Design?
Compound Design is Revilico's interactive molecular structure editor, enabling medicinal chemists and computational scientists to draw, edit, and export chemical structures directly within the platform. New compound ideas can be sketched from scratch, existing structures modified for analog design, and the resulting SMILES exported immediately to docking, property prediction, or generative chemistry engines without leaving the Revilico environment.
## Background
Structure drawing is a core activity in drug discovery, from initial hit design through lead optimization. Having a drawing tool integrated directly with computational engines eliminates the friction of exporting from a standalone editor, converting file formats, and re-uploading. Compound Design uses the Ketcher chemical structure editor, a browser-based drawing tool that produces valid SMILES, MOL, and SDF output and supports the full range of organic chemistry including stereochemistry, isotope labels, R-groups, and query atoms.
## Interface
**Left toolbar (drawing tools):**
* Bond drawing tools: single, double, triple bonds and aromatic ring
* Atom selection and movement tools
* Text and label tools
* Stereo bond tools (wedge, dash)
* R-group and S-group tools
* Reaction arrow tools
* Chain and template tools
**Right atom palette:**
The quick-access atom palette on the right side of the canvas provides single-click placement for the most common elements in medicinal chemistry: H, C, N, O, S, P, F, Cl, Br, I.
**Bottom ring template bar:**
Pre-built ring templates are available at the bottom of the canvas for common ring systems: cyclopentane, cyclohexane, benzene, cyclopentadiene, cycloheptane, cyclooctane, and fused ring starters.
**Top toolbar:**
* File operations: new, open, save, copy, paste, delete
* Undo and redo (also available via Ctrl+Z / Ctrl+Y)
* Zoom in and out
* Analysis tools: calculate properties, check structure
* 3D view toggle
* Atom map and R-group label tools
* Mode selector (Molecules / Reactions)
## Workflow
1. Draw a structure using the bond and atom tools, or open an existing structure file
2. Edit substituents, ring systems, stereochemistry, and functional groups using the drawing tools
3. Use the mode selector to switch between molecule drawing and reaction scheme drawing
4. Export the structure as SMILES for use in any Revilico engine input, or save as MOL/SDF for downstream file-based workflows
## Running the Engine
### Inputs
| Input | Description |
| ------------- | ------------------------------------------------------------------- |
| New structure | Draw from scratch using the on-canvas tools |
| Existing file | Open a MOL, SDF, or SMILES file from local disk or Data Engineering |
### Outputs
* **SMILES:** Canonical SMILES string for direct copy-paste into any Revilico engine
* **MOL/SDF file:** 2D structure file for download or export to file-based workflows
* **Reaction SMILES:** For reaction scheme mode, outputs a reaction SMILES string
# Compound Review
Source: https://docs.revilico.bio/docs/revcompound-review
Collaborative Compound Review Sessions for Molecular Libraries with SMILES-Based Annotation and Shared Analysis
## Why Use Compound Review?
Compound Review is Revilico's collaborative workspace for reviewing, annotating, and discussing molecular libraries across a research team. Instead of passing spreadsheets over email or maintaining separate local copies, teams create named review sessions from a shared SMILES dataset and work through molecular annotations, computed property analyses, and go/no-go decisions within a single tracked environment.
## Background
Lead optimization and compound triage in drug discovery require iterative review cycles where medicinal chemists, computational scientists, and project managers evaluate the same set of molecules from different angles. Compound Review centralizes this process by attaching a review session to a molecular dataset, enabling annotations, analysis results, and shared notes to accumulate in one place rather than fragmenting across individual workflows.
## Creating a Review Session
A review session is created by providing a session name (e.g., "Lead Optimization Batch 1"), an optional description for project context, and uploading a CSV file containing molecular SMILES data. The SMILES column is auto-detected from common header names including `smiles`, `smileString`, `SMILES`, and similar variants, and can appear at any position in the file. Molecule IDs can be auto-generated from the row index if the input file does not contain an explicit identifier column.
Once created, the session is accessible under the Compound Review tab for the creating user and under the Shared with Me tab for all team members with access.
## Interface Tabs
**Compound Review tab** — The primary workspace. Displays the molecular library loaded from the uploaded CSV with 2D structure rendering for each compound. Reviewers can annotate individual molecules, flag compounds for follow-up, and record decisions.
**Analysis tab** — Connects to Revilico's computational engines to run property predictions, similarity searches, or docking scores against the loaded compound set and display the results in the context of the review session.
**Shared with Me tab** — Lists all review sessions shared by other team members, enabling cross-functional access to the same dataset and its accumulated annotations without requiring file transfers.
## Running the Engine
### Inputs
| Field | Required | Description |
| -------------------------- | -------- | -------------------------------------------------------------------- |
| Review Name | Yes | Descriptive name for the session (e.g., "Lead Optimization Batch 1") |
| Description | No | Free-text notes about the purpose or context of the review |
| Auto-generate Molecule IDs | No | Assign sequential IDs when the input file lacks an identifier column |
| CSV File | Yes | CSV with a SMILES column; header variants auto-detected |
### Outputs
* **Compound library view:** 2D structure rendering for each molecule in the uploaded set
* **Annotation layer:** Per-molecule flags, notes, and decision records
* **Analysis integration:** Computed properties and engine outputs attached to the session
* **Shared access:** Session visible to all invited team members under their Shared with Me tab
# RevConformer
Source: https://docs.revilico.bio/docs/revconformer
Running a Conformer Search Screening for a High Throughput Set of Molecules
## Why Use this product?
The Conformer search engine is a computational tool that delivers thermally accessible conformation ensembles that capture the structural flexibility of drug-like molecules essential for accurately binding predictions and structure activity relationship analysis. You will use this engine when you need to generate comprehensive ensembles of low-energy 3D molecular conformations for drug discovery applications like molecular docking or pharmacophore modeling. You can utilize this engine to analyze intrinsic compound conformational energies when cross referencing generated/docking ligand protein poses. Large deviations in conformation needs to be analyzed to ensure favorable energetic compositions and changes.

## Background
Conformation refers to the specific 3D shape a molecule can adopt. We may have a single molecule, however this molecule can adopt several different plausible shapes. But why is this important? The conformation of a molecule affects how effectively it binds to targets, influencing potency, absorption, and side effects. Because conformations have different energies, finding lower energy conformations allows researchers to better predict drug-target interactions, and design molecules for optimal binding. When discovering different conformations, chemists often discover them using advanced spectroscopy, X-ray crystallography, and electron microscopy. This can be considered a very time consuming and expensive task. This is where Revilico’s Conformer Search Engine comes into place. This engine operates as a high throughput conformer predictor that can both effectively and speedily discover many different conformations of your small molecule, saving both the time and money to run any experimentation.
Now how does it work? We first start with our SMILES strings. We can convert this into a 3D structure Before introducing the ETKDG v2 algorithm. This algorithm was designed based on experimental conformational data where statistics regarding typical bond lengths, typical angles, and preferred torsion angles were hard coded into this algorithm. This algorithm will take the base molecule and generate different possible confirmations ensuring that bonds and non-bonded atoms fall within a realistic distance range, certain rotations are sampled more frequently than others, and rings are constrained to known geometries. The step by step process of generating a single conformer is first to build a distance bounds matrix, followed by sampling torsion angles based on experimental data, solving the geometry problem ensuring that atoms are placed in a manner that satisfies the constraints, and finally the algorithm checks for any possible steric clashes. We will now generate a large panel of different conformations for each molecule that was inputted to get a probability landscape of that conformer’s existence.
We now introduce the MMFF94 forcefield. In this step we will compute the force-field energy, slightly move the atoms to create an energetic downhill motion towards an energetic minimum, and stop when the forces are small, relieving strain, fixing small clashes and producing a local minimum conformer. This forcefield is based on these equations that assess internal compound interactions:
$$
E = \sum \text{interactions}
$$
$$
E_{\text{total}} = E_{\text{bonds}} + E_{\text{angle}} + E_{\text{dihedral}} + E_{\text{VDW}} + E_{\text{elec}}
$$
$$
E_{\text{bond}} = \sum k_b (r - r_0)^2
$$
$$
E_{\text{angle}} = \sum k_0 (\theta - \theta_0)^2
$$
$$
E_{\text{dihedral}} = \sum \frac{V_n}{2} (1 + \cos(n\phi - \gamma))
$$
$$
E_{\text{VDW}} = \sum 4\epsilon \left[ \left( \frac{\sigma}{r} \right)^{12} - \left( \frac{\sigma}{r} \right)^6 \right]
$$
$$
E_{\text{elec}} = \sum (q_i \times q_j)/(4\pi \epsilon_0 r_{ij})
$$
This equation of total energy will calculate the total energy results derived from the use of the forcefield. When we want to move the atoms downhill to get to a local minimum conformer, this will be based on calculating the gradient of the forcefield/ energetic landscape as denoted by:
$$
F = - \nabla E
$$
Where ∇E calculates the gradient derivate of the energy in each possible direction, in our case along the x, y and z plane. This can be denoted as
$$
\nabla E(r) = \left( \frac{\partial E}{\partial x}, \frac{\partial E}{\partial y}, \frac{\partial E}{\partial z} \right)
$$
We will do this for each conformer that was generated to produce a list of different conformers and their associated energies. To curate the panel of conformers we will return, we will first filter all conformers based on their energy, with lowest energy being highly prioritized and highest energy being listed last. We will then down-select our panel based on our energy window. Our energy window is the energy difference between the target conformer generated and the conformer with the lowest energy. This can be denoted as a
$$
\nabla E = E_{\text{conformer}} - E_{\text{min}}
$$
Where ∇E must be less than the size of the energy window. If it does not satisfy this requirement we will get rid of the conformer from #down-selected list. The next step will be to filter based on the Root Mean Square Deviation (RMSD) thresholds, which ensures proper geometric feasibility and accuracy. For this we do not want conformers that are too similar to each other, or in our case less than the RMSD difference threshold measured between each conformer generated. What we will do is evaluate each pair of conformers, removing the conformer with the higher energy. In the end we will only return conformers that are greater than our RMSD threshold (to ensure conformer diversity), and that perform better, with lower calculated energies.
## Interactive Results Viewer
Explore conformer search results interactively. View 3D conformer structures with an energy slider, Boltzmann populations, and shape descriptors.
# Cytokine Release
Source: https://docs.revilico.bio/docs/revcytokine-release
Drug-Induced Cytokine Release and Cytokine Storm Risk Prediction Across Cell Lines
## Why Use This Engine?
In the documentation below, we will use Revilico's Cytokine Release assay engine to predict the inflammatory cytokine response to drug treatment across a cancer cell line panel. This assay identifies drugs at concentrations that induce excessive cytokine secretion, flagging cytokine storm risk before experimental testing. Cytokine storm is a potentially fatal immune cascade associated with high-dose cytotoxic therapies and certain immunomodulatory agents.
## Background
Cytokines are small signaling proteins secreted by cells in response to stress, damage, or immune activation. When a drug induces significant cellular stress or direct immune activation, cytokine levels in the cell culture supernatant rise in proportion to the magnitude of the stress response. At very high drug concentrations or in immunocompetent contexts, cytokine levels can reach thresholds associated with systemic inflammatory toxicity.
The Cytokine Release module models six core cytokines: IL-6, IL-8, IL-1 beta, TNF-alpha, IFN-gamma, and IL-10. The database encompasses over 60 catalogued cytokines spanning the interleukin, chemokine, interferon, TNF superfamily, colony-stimulating factor, and growth factor families. Cytokine levels are derived from the cellular stress signal generated by the underlying GNN+MLP viability model.
## Simulation Model
For each cytokine, a baseline secretion level at zero drug effect and a maximum fold-induction at saturating stress are defined from literature-curated values for typical cancer cell lines:
| Cytokine | Baseline | Max fold-induction |
| --------- | --------- | ------------------ |
| IL-6 | 45 pg/mL | 22x |
| IL-8 | 180 pg/mL | 28x |
| IL-1 beta | 8 pg/mL | 18x |
| TNF-alpha | 28 pg/mL | 20x |
The cytokine level at stress $s = 1 - v(c)$ is computed as:
$$
[\text{Cytokine}](c) = B_k \cdot \left(1 + s \cdot (F_k - 1)\right) \cdot \eta, \quad \eta \sim \mathcal{LN}(0, 0.18^2)
$$
Where $B_k$ is the baseline for cytokine $k$, $F_k$ is the maximum fold-induction, and $\eta$ is a lognormal multiplicative noise term with 18% coefficient of variation reflecting inter-well variability.
**Cytokine Storm Flag**
A cytokine storm event is flagged when any cytokine exceeds a per-cytokine threshold (default 500 pg/mL):
$$
\text{Storm} = \bigvee_k \left[ [\text{Cytokine}_k](c) > \theta_k \right]
$$
The storm flag is reported per concentration per cell line, enabling identification of the minimum storm-inducing concentration. At clinical exposure concentrations (typically 1 to 5 micromolar for most kinase inhibitors), storm risk is generally absent. Storm flags at extreme concentrations (100 micromolar) reflect the acute toxicity profile rather than the therapeutic window.
## Parameters
| Parameter | Default | Description |
| ---------------------- | -------------------------------------- | -------------------------------------------- |
| Selected cytokines | IL-6, IL-8, IL-1b, TNF-a, IFN-g, IL-10 | Which cytokines to display |
| Storm threshold | 500 pg/mL | Per-cytokine level triggering the storm flag |
| Exposure hours | 72 h | Duration of drug treatment |
| Hill coefficient | 1.2 | Dose-response steepness |
| Biological variability | None | Noise level for replicates |
## Outputs
* **Cytokine levels (pg/mL):** Per-cytokine concentration at each drug dose for each cell line
* **Fold-induction:** Ratio of drug-treated to untreated secretion per cytokine
* **Cytokine storm flag:** Binary indicator per concentration per cell line with minimum storm dose
* **Multi-cytokine heatmap:** All cytokines by concentration, colored by level relative to threshold
* **Dose-response curves:** Per-cytokine secretion curves across the dose range
* **96-well heatmap:** Plate-view colored by selected cytokine level or storm risk
* **Cytokine family breakdown:** Grouped view across interleukins, chemokines, interferons, and other families
# RevDenovo
Source: https://docs.revilico.bio/docs/revdenovo
Generating a Compound Library Using Revilico’s De Novo Library Generator
## Why Use this product?
Revilico’s De Novo Library Generation Engine is an AI powered tool that leverages reinforcement learning and pretrained generative models and scoring functions to design drug-like molecules, perform scaffold hopping, connect molecular fragments with linters, and generate peptides with natural and non natural amino acids. This Engine will enable rapid exploration of chemical space for hit discovery, lead optimization, and library enumeration at scales from hundreds to millions of compounds. This library generated will eventually serve as a concentrated screening set for other engines on the platform to narrow down synthesizable candidates.
## Background
Oftentimes in drug discovery, we already have a target protein and would like to screen a compound library to see which molecules have the best binding affinity, and are most promising to move to downstream analysis such as Molecular Dynamics and Free Energy Perturbation Calculations. This oftentimes requires the user to already have a molecular library built out in which they would like to screen, oftentimes a library of compounds that have some sort of desired property, may be on hand, or are easily synthesizable or accessible because the lab next door has them on hand. But what if we do not have a library and only a target molecule? We then can use Revilico’s De Novo Library Generation engine which is a core capability in AI-assisted drug discovery that attempts to solve the inverse design problem: Given a desired set of properties, create new compounds likely to satisfy them. Rather than enumerating combinations of fragments or scaffolds, De Novo Generation builds molecules from scratch using learned language or graph models that encode valid chemical structure syntax. This is an alternative to pre-determined library selection and allows for less molecules to be synthesized and tested to get towards candidate hits because we are ‘creating a new needle’ rather than finding a needle in a haystack.
Now how does it work? The basis of this engine is a chemical language model or neural network that treats SMILES strings like sentences and treats atoms and bonds like tokens or words. This model has been trained on a large molecular database with valid, synthesized small molecules. From this model it has learned valence rules, common ring systems, typical functional groups, and what molecules look like medicinal chemistry vs which look like chemistry. Essentially this model is sampling from learned chemical patterns.
We will first start with a start token, where a token is a small piece of a SMILES string. The job of the neural network is to predict the probability distribution of the next possible token. The next token is not always the highest-probability token, since that would lead to the same molecule being generated over and over again. Instead we sample from the probability distribution (or the latent space vector representation of the molecules). This sample is then appended to the current token or current SMILES string. We will repeat this process of generating the distribution, sampling from the distribution, and appending until we have generated the end token, signaling the end of one molecule. We will do this for several molecules generating a library of valid chemical candidates.
Now we have a set of candidate molecules, all of which are chemically and synthetically valid, however some may be undesirable molecules or molecules are impractical. We are now asking the question which of the molecules should we keep, and which do we remove? To do this we go through a series of different quality control check, first being a validity check to see if the SMILES string is synthetically valid, whether it can be parsed into a molecular graph, whether it contains allows atom types, and whether the molecule is within a maximum size limit. The next step is a chemical sanity filter that will be constrained to a molecular weight range, number of rings,and a number of heavy atoms. We will then lastly go through a diversity handling step, where we will remove duplicates, as well as conduct a similarity based pruning. This is to ensure we have wide coverage of the entire chemical space.
To run the De Novo Library Generation Pipeline: (1) Select your molecule generation type from five options: De Novo Generation (no input molecules), Scaffold Decoration (scaffold-based generation), Fragment Linker Design (link molecules with linkers), Molecular Optimization (lead-based generation), or Peptide Generation (peptide design with natural and non-natural amino acids). (2) Configure the Pipeline Name and select your model, defaults are included and don’t need to be changed, with Molecular Optimization offering seven model options (Low similarity, Medium similarity, High similarity, Real medicinal chemistry transformation, Scaffold-based transformations, Broad scaffold diversification, or custom model upload). (3) Upload or submit input SMILES if required by your selected generation type, adjust molecular diversity settings between Exploratory and Focused (with Temperature slider ranging from "Conservative" to "Wild" for Exploratory mode), and enter the number of output molecules desired per input. (4) Press "Create Pipeline" to initiate generation, then retrieve results in the Analysis section once completed, outputs include generated molecular structures with diversity scores and property predictions.
# RevDrive
Source: https://docs.revilico.bio/docs/revdrive
Centralized File Storage, Organization, and Sharing for All Revilico Research Assets
## Why Use RevDrive?
RevDrive is Revilico's central file management system, providing a shared workspace where all molecular files, datasets, computational results, and research assets are stored, organized, and accessible to the team. Rather than maintaining files on local machines or in disconnected cloud storage, RevDrive keeps all research data in one place with team sharing, versioned organization, and direct integration with every Revilico engine for drag-and-drop input.
## Interface
**Left navigation:**
* **My Files:** Personal file storage accessible only to the current user
* **Shared Files:** Files explicitly shared with the user by team members
* **SMILES to CSV:** Built-in conversion utility to generate a CSV file from a list of SMILES strings, ready for use in any engine that accepts tabular molecular input
* **Shared Drives:** Team-level drives with shared access across the organization
* **Starred:** Files and folders marked as favorites for quick access
* **Recent:** Files accessed or modified most recently
* **Trash:** Deleted files pending permanent removal
**Main file browser:**
Files are displayed in a sortable grid or list view. The toolbar provides:
* **Search:** Full-text search across file names and metadata
* **Project filter:** Scope the view to a specific project
* **Type filter:** Filter by file type (structure, trajectory, dataset, etc.)
* **Modified sort:** Sort by last modification date
* **Upload:** Upload files from local disk
* **Grid/List toggle:** Switch between visual grid and compact list layouts
Each file entry shows the filename, last modified date, file size, and an action menu for rename, move, share, download, and delete operations.
**Storage indicator:** Total storage used is displayed at the bottom of the left navigation panel.
## SMILES to CSV
The built-in SMILES to CSV utility accepts a list of SMILES strings (one per line or comma-separated) and converts them to a properly formatted CSV file with auto-generated molecule IDs. The resulting file can be saved to RevDrive and used directly in Compound Review, RevViability, RevAssay, or any engine that accepts a CSV input with a SMILES column.
## Running the Engine
### Inputs
| Action | Description |
| ------------- | ------------------------------------------------------------------ |
| Upload | Upload any file type from local disk to My Files or a Shared Drive |
| SMILES to CSV | Paste SMILES strings for batch conversion to a structured CSV |
| Create folder | Organize files into named project folders |
### Outputs
* **Stored files:** All uploaded files persisted and accessible across sessions
* **Shared files:** Files made available to named team members or shared drives
* **CSV export:** Structured CSV from the SMILES to CSV converter, ready for engine input
* **Downloaded files:** Any file available for local download at any time
# RevEdit
Source: https://docs.revilico.bio/docs/revedit
Quick Edit — Spreadsheet Editor and Molecular Viewer for Research Data Files
## Why Use RevEdit?
RevEdit is Revilico's in-browser file editor, combining a full spreadsheet editing environment with an integrated molecular viewer in a single interface. Researchers can open tabular data files (CSV, Excel, ODS) alongside molecular structure files and trajectories, making edits and visualizing the corresponding 3D structures without switching between separate applications. This is particularly useful for reviewing docking results tables alongside binding poses, editing compound libraries while viewing structures, or annotating MD trajectory data while watching the simulation.
## Interface
**Left panel (file and shortcuts):**
The Current File section displays the active file with its format badge. The Change File button opens a file picker to switch to a different file from RevDrive or local disk.
Keyboard shortcuts are listed for quick reference:
* **Ctrl+S:** Save
* **Ctrl+Z:** Undo
* **Ctrl+Y:** Redo
Two viewer modes are shown as selectable cards:
**Data Editor** supports tabular file formats: CSV, TSV, XLS, XLSX, ODS. This mode opens the file in the spreadsheet editor panel for row-level editing, column management, and data export.
**Molecular Viewer** supports a comprehensive range of molecular file types:
* **Structures:** PDB, CIF, MMCIF, BCIF, SDF, MOL, XYZ, GRO
* **Trajectories:** DCD, XTC, TRR, PDBQT
* **Density maps:** CCP4, MAP, MRC, DSN6
* **Topology:** PSF, TOP, ITP, PRMTOP
**Right panel (Spreadsheet Editor):**
The spreadsheet editor displays the file contents as an editable grid. The toolbar provides:
* **Save:** Save the current file back to RevDrive (Ctrl+S)
* **Export:** Download the file in CSV, TSV, XLSX, or other formats
* **Add:** Insert new rows or columns
* **Delete Row:** Remove the selected row
* **Columns:** Show, hide, or reorder columns
* **Search:** Filter rows by keyword across all columns
* **Undo/Redo:** Step through edit history
Cells are edited by clicking. Row numbers can be clicked to select entire rows for deletion or batch operations. The row count is shown in the top right corner.
## Workflow
1. Open a file from RevDrive or upload from local disk
2. Select Data Editor for tabular files or Molecular Viewer for structure files
3. Make edits in the spreadsheet or adjust the 3D view
4. Save changes back to RevDrive or export to a new format
## Running the Engine
### Inputs
| Format type | Supported formats |
| ----------- | ----------------------------------------- |
| Spreadsheet | CSV, TSV, XLS, XLSX, ODS |
| Structure | PDB, CIF, MMCIF, BCIF, SDF, MOL, XYZ, GRO |
| Trajectory | DCD, XTC, TRR, PDBQT |
| Density | CCP4, MAP, MRC, DSN6 |
| Topology | PSF, TOP, ITP, PRMTOP |
### Outputs
* **Saved file:** Edited file written back to RevDrive
* **Exported file:** Downloaded in any supported output format
* **3D visualization:** Interactive molecular view of the loaded structure or trajectory
# RevFEP
Source: https://docs.revilico.bio/docs/revfep
Cloud-Native Free Energy Perturbation for Binding Affinity Prediction Powered by OpenFE, OpenMM, and AWS Batch
## Why Use This Engine?
In the documentation below, we will use Revilico's RevFEP engine to compute binding free energies between protein targets and small molecule ligands using alchemical free energy perturbation (FEP). RevFEP automates the full pipeline from raw protein-ligand complex PDB files through structure preprocessing, alchemical network construction, GPU-accelerated simulation, and statistical analysis to produce DDG or DG estimates with uncertainty quantification. It is the highest-accuracy binding affinity method available on the platform and is the appropriate engine when docking scores are insufficient to distinguish closely related lead compounds.
## Background
Free energy perturbation is a rigorous thermodynamics-based method for computing binding affinity differences between related molecules. Rather than using empirical scoring functions, FEP simulates the alchemical transformation between two chemical states using molecular dynamics and computes the free energy difference from the work distribution of the transformation. When applied to a series of ligands against a protein target, FEP produces binding affinity rankings that are far more accurate than docking or MM-GBSA, with mean absolute errors typically below 1 kcal/mol for well-prepared systems.
RevFEP is built on OpenFE 1.10 and OpenMM, deployed on AWS Batch for scalable GPU execution. It supports four calculation types covering the range from rapid relative ranking to absolute binding free energy estimation, and implements seven advanced protocol stages including non-equilibrium MD, REST2 enhanced sampling, adaptive lambda optimization, and MM-GBSA pre-filtering.
## Calculation Types
**RBFE (Relative Binding Free Energy)**
RBFE calculates the relative binding affinity between pairs of ligands (DDG\_bind). It is the most efficient engine for lead optimization where a series of structural analogues must be ranked against a known binder. Pairs of ligands are connected in a perturbation network, and for each pair a single alchemical transformation is simulated. The OpenFE HREX/MBAR protocol is used with Hamiltonian replica exchange across lambda windows.
$$
\Delta\Delta G_{\text{bind}} = \Delta G_{\text{bind}}^{\text{ligand B}} - \Delta G_{\text{bind}}^{\text{ligand A}}
$$
**ABFE (Absolute Binding Free Energy)**
ABFE calculates the absolute binding free energy of a single ligand to a protein (DG\_bind) by running two independent simulation legs: the complex leg (protein and ligand together) and the solvent leg (ligand in water only). ABFE is slower than RBFE but provides absolute DG values without requiring a reference compound.
$$
\Delta G_{\text{bind}} = \Delta G_{\text{complex}} - \Delta G_{\text{solvent}}
$$
**SepTop (Separated Topologies)**
SepTop uses separated topology alchemical transformations where the two end-state ligands are represented as independent topology files simultaneously scaled between lambda = 0 and lambda = 1. This allows perturbations between structurally dissimilar ligands where a maximum common substructure mapping would be too small for reliable RBFE.
**AHFE (Absolute Hydration Free Energy)**
AHFE calculates the free energy of transferring a ligand from vacuum into water (DG\_hydration). It is a protein-free calculation used to validate force field parameters and assess ligand solvation characteristics before running more expensive binding calculations.
## Structure Preparation Pipeline
**Protein Preparation**
The input protein-ligand complex PDB is parsed to separate protein ATOM records from ligand HETATM records. The protein cleaning pipeline then runs in sequence: standard residues are retained while non-standard residues and crystallographic artifacts are handled; missing heavy atoms are rebuilt using PDBFixer; missing hydrogen atoms are added at pH 7.0; disulfide bonds and CONECT records are preserved; and the cleaned structure is written as `preprocessed/protein.pdb` for use as the ProteinComponent in all subsequent calculations.
**Ligand Extraction and Processing**
Ligands are extracted from the input PDB using a two-tier strategy that handles both GNINA docking output (HETATM ligand with CONECT bond order records) and generic docking output (ATOM ligand with UNL residue name). The extracted ligand is passed through RDKit for hydrogen addition, 3D coordinate assignment, and SMILES round-trip validation to ensure chemical correctness.
**3D Alignment for RBFE and SepTop**
For multi-ligand campaigns, ligands extracted from independent docking runs may be oriented differently in the binding pocket. The engine performs Maximum Common Substructure (MCS) alignment using RDKit `rdFMCS` with the first ligand as the reference, applying `rdMolAlign.AlignMol()` with the MCS atom map for 3D superposition. Ligands with an MCS smaller than 3 atoms are retained with a warning but are not aligned.
**Partial Charge Assignment**
Partial charges are pre-assigned during preprocessing using the selected small molecule force field (OpenFF Sage or Rosemary by default, GAFF2 available as an alternative). Pre-assigning charges avoids running antechamber at simulation time and ensures consistent charge treatment across all ligands in the campaign.
## Perturbation Network
For RBFE and SepTop, ligands are connected in a perturbation network that defines which pairs will be simulated. The network topology determines the total number of alchemical transformations and therefore the size of the AWS Batch array job. The engine constructs the network using Kartograf atom mapping, which relies on 3D spatial overlap of aligned ligand poses to identify the maximum common substructure for each pair.
## Three-Phase AWS Pipeline
**Phase 1: Setup**
All inputs are preprocessed, the perturbation network is constructed, and OpenFE transformation JSON objects are written to `openfe_setup/transformations/`. A manifest file lists all transformations with their array indices for the run phase.
**Phase 2: Run (GPU Array Job)**
Each AWS Batch array element runs one alchemical transformation. The transformation is loaded from its JSON, the lambda schedule is applied, and HREX molecular dynamics is executed across all lambda windows on a GPU. Each window produces a `dhdl.xvg` file. The GPU array job runs all transformations in parallel, making the total wall time proportional to the slowest individual transformation rather than the sum.
**Phase 3: Gather and Analysis**
After all array elements complete, the gather phase collects per-transformation result JSONs and runs MBAR (Multistate Bennett Acceptance Ratio) analysis via alchemlyb to produce the final DG or DDG estimates with uncertainty. Results are written to `openfe_summary/summary.tsv` and `analysis_output/` containing MBAR convergence plots and the DDG matrix.
## Advanced Protocol Stages
**Stage 1: Advanced Force Fields**
The protein force field can be set to AMBER ff14SB (default) or ff19SB. The small molecule force field can be set to OpenFF Sage 2.2 (default), OpenFF Rosemary, or GAFF2. Water model options include TIP3P, TIP4P-Ew, and SPC/E.
**Stage 2: Non-Equilibrium MD (NE-MD)**
NE-MD implements the Crooks Fluctuation Theorem using the Bennett Acceptance Ratio estimator for rapid DG estimation. Rather than running long equilibrium simulations at multiple lambda windows, NE-MD switches the Hamiltonian rapidly between two end states and measures the forward and reverse work distributions. BAR analysis of the bidirectional work values yields a free energy estimate:
$$
\Delta G = k_BT \ln \frac{\langle f(W_{\text{rev}} + \Delta G) \rangle_{\text{rev}}}{\langle f(W_{\text{fwd}} - \Delta G) \rangle_{\text{fwd}}}
$$
Where $f(x) = 1/(1 + e^{x/k_BT})$ is the Fermi function and the averages are over forward and reverse work trajectories.
**Stage 3: REST2 Enhanced Sampling**
Replica Exchange with Solute Tempering version 2 (REST2) improves conformational sampling by selectively scaling the ligand (solute) interaction energies across replicas at effective temperatures higher than the target. This allows the ligand to overcome torsional barriers that would not be crossed in standard equilibrium MD, improving convergence without the cost of scaling the full system.
**Stage 4: Adaptive Lambda Windows**
After a pilot FEP run, this stage analyzes the MBAR phase-space overlap between adjacent lambda windows. Windows with overlap below the target threshold are identified, and additional intermediate windows are suggested. The refined lambda schedule is written to `adaptive_lambda/lambda_suggestions.json` for use in the production run.
**Stage 5: Force Field Torsion Scan**
Dihedral energy scans are run across all rotatable bonds in each ligand using multiple force field methods (OpenFF, GAFF2, and semiempirical xTB). Comparing energy profiles across methods identifies potential force field parameter quality issues before committing to expensive FEP simulations.
**Stage 6: Tighter ABFE Settings**
Increases the lambda window count and tightens the restraint configuration for ABFE calculations, improving accuracy at the cost of longer simulation time. Recommended for final-stage ABFE calculations on clinical candidates.
**Stage 7: MM-GBSA Pre-Filter**
Before committing to FEP, all ligands are rapidly scored against the protein using MM-GBSA with Generalized Born (OBC2) implicit solvent:
$$
\Delta G_{\text{MM-GBSA}} \approx E_{\text{complex}}^{\text{GB}} - E_{\text{protein}}^{\text{GB}} - E_{\text{ligand}}^{\text{GB}}
$$
All three energies use the same OBC2 solvation model (solvent dielectric 78.5, solute dielectric 1.0) with local minimization and short equilibration, so systematic errors cancel in the difference. Ligands below the score cutoff are pruned from the perturbation network, reducing the total number of FEP transformations and the associated compute cost.
## Running the Engine
### Inputs
| Parameter | Default | Description |
| -------------------------- | --------------- | ---------------------------------------------------------- |
| Protein-ligand complex PDB | Required | One PDB per ligand (RBFE/SepTop) or single PDB (ABFE/AHFE) |
| Calculation type | rbfe | `rbfe`, `abfe`, `septop`, or `ahfe` |
| Protein force field | ff14SB | `ff14SB` or `ff19SB` |
| Small molecule force field | OpenFF Sage 2.2 | `openff-2.2`, `openff-rosemary`, or `gaff2` |
| Water model | TIP3P | `tip3p`, `tip4pew`, or `spce` |
| Lambda windows | 11 | Number of alchemical intermediate windows |
| Advanced stages | None | Any combination of the 7 enhancement stages |
| MM-GBSA cutoff | -20 kcal/mol | Score threshold for pre-filter pruning (Stage 7) |
### Outputs
* **DDG matrix:** Relative binding free energy differences between all ligand pairs (RBFE/SepTop)
* **DG table:** Absolute binding free energies per ligand (ABFE/AHFE)
* **Uncertainty estimates:** MBAR standard error per transformation
* **MBAR convergence plots:** Phase-space overlap matrices per transformation
* **Perturbation network visualization:** Graph of all simulated ligand pairs
* **MM-GBSA scores:** Pre-filter scores and pruned network (Stage 7)
* **Adaptive lambda report:** Suggested refined lambda schedule (Stage 4)
* **NE-MD work distributions:** Forward and reverse work histograms with BAR estimates (Stage 2)
* **Torsion scan profiles:** Per-bond dihedral energy plots across force fields (Stage 5)
# RevGeometry
Source: https://docs.revilico.bio/docs/revgeometry
Understanding Interaction Dynamics using Revilico’s Pharmacophore Analysis Engine
## Why Use this product?
Revilico’s Geometry Minimization and Thermochem Engine rapidly minimizes molecular geometries and computes temperature-dependent thermochemical properties, enabling fast prediction of reaction energetics, conformational stability, and chemical feasibility without expensive quantum mechanical calculations. This engine is best used when you need to optimize molecular structures to their lowest-energy conformations and predict thermodynamic properties for understanding reaction thermodynamics and molecular stability. This engine is key to understanding the driving energetics of intrinsic compound conformations to compare against protein ligand interaction conformations. This elucidates the energetic penalties of shifting compound geometries in pockets of proteins to ensure that thermodynamically, your system makes sense.

## Background
How does this engine work? We first start with our SMILES string input which we then convert into a 3D molecular structure. We then feed this structure into a model that will predict the energy as well as the energy gradient, giving the engine an instruction for which direction to move atoms in order to reach the lowest possible energy state for that molecule. By slowly adjusting the coordinates and positioning of the atoms in space and continuously calculating energies, we can find local thermodynamic equilibriums.
Here is how the model works. Typically when calculating the energetics of a molecule we will use force fields, where we will calculate the total bond energy by summing up energy from the bonds, angles, torsions, van der Waals, steric hinderance, and electrostatics. Traditional forcefields are fast, but they are more limited in accuracy. The more accurate way of doing this is by doing large scale Density Functional Theory (DFT) calculations, however doing this at scale is very computationally expensive. Traditionally, DFT has been used to help approximate the physical interactions and energetics of electrons within biochemical systems, and is a computationally feasible way to find approximate solutions to Schrodinger’s equations without unimaginable amounts of computation. DFT usually follows a general framework based on the following equation:
$$
\left[ -\frac{1}{2} \nabla^2 + V_{\text{ext}}(r) + V_H(r) + V_{XC}(r) \right] \psi_i(r) = \epsilon_i \psi_i(r)
$$
Where ψ(r) is Kohn\_Shams orbital, -½ ∇² is the kinetic energy, V\_ext(r) is the external potential, V\_H(r) is Hartree term, and V\_XC(r) is the exchange correlation. DFT has been a widely used solution to analyzing chemical systems at the quantum level, but recent breakthroughs and data collections with UMA/QMOL25 datasets by Fairchem also enable for the training of AI models on these datasets for increased speed and accuracy.
This engine introduces a machine learned interatomics potential (MLIP) trained on large scale Density Functional Theory data (DFT) from the Fairchem team at Meta. This data comes from a quantum mechanical calculation that will model electron density, electronic energy, and nuclear forces at scale across chemical and protein systems. The data itself contains atomic numbers, 3D atomic coordinates, total electronic energy from DFT calculations, and atomic forces or the gradient of the DFT energy. The data will often contain multiple geometries of the same molecule so it is trained on many off-equilibrium geometries so it can learn how energy changes when atoms move. How the model works is that for each molecule, the model will receive the atomic identity, atomic position, and their local atomic environment (constructing a local neighborhood around each atom to see which atoms are nearby, their distances, and their relative orientation). The model will output the total energy which is an approximation to the DFT total electronic energy, and the negative gradient of the energy, which will show directionality for each atom to reach its lowest energy state and its eventual geometrically optimal position in 3D space with minimized energetics.
At this point, the molecule is at a stable, low energy 3D conformation. We want to compute our thermochemistry step. Thermochemistry can be defined as studying how the heat energy changes during chemical reaction, with an emphasis on how the heat is absorbed or released. Essentially what we want to understand is how stable a molecule is, how it responds to temperature, and how favorable processes involving this molecule are. Thermochemistry is based on this equation
$$
G = H - TS
$$
Where G is the Gibbs free energy or the balance between energy and disorder, H is enthalpy or how much internal energy the molecule has, S is entropy and how much freedom or disorder the molecule has and T is temperature.
For measuring thermochemistry, the engine will capture the following properties. First, atomic masses, where heavier atoms move more slowly meaning lower vibrational frequencies. Second, molecular geometry, where this will define how atoms are connected and how rigid or flexible the structure is. Third, moments of inertia, which is derived from atomic masses and is their distance from the center of mass. This is required to compute rotational entropy. Fourth, potential energy surface curvature, where it will answer the question if we move the atoms away from this geometry, how quickly does the energy increase. This comes from the second derivatives of energy and is estimated using the MLIP from DFT calculations. Here steep curvatures mean stiff vibrations and low entropy. Lastly, molecular symmetry, where symmetry will affect rotational entropy.
The main goal of this engine is to compute enthalpy and entropy. To calculate this we must understand how the molecules vibrate. To do this the engine will slightly perturb the atom positions. It will observe how the forces change, building a matrix describing how motion couples between atoms. Then it will diagonalize the matrix.
Enthalpy is calculated with the following equation:
$$
H = H_{\text{elec}} + H_{\text{trans}} + H_{\text{rot}} + H_{\text{vib}}
$$
Where:
$$
H_{\text{elec}} \approx E_{\text{elec}}
$$
$$
H_{\text{trans}} = \frac{3}{2} kT
$$
$$
H_{\text{rot}} = \frac{3}{2} kT
$$
$$
H_{\text{vib}} = \sum_i \left[ \frac{h\nu_i}{2} + \frac{h\nu_i}{\exp(h\nu_i/kT) - 1} \right]
$$
Entropy is calculated with the following equation:
$$
S = S_{\text{trans}} + S_{\text{rot}} + S_{\text{vib}}
$$
Where:
$$
S_{\text{trans}} = k \left[ \ln \left( \frac{(2\pi mkT)^{3/2}}{h^3} \frac{V}{N} \right) + \frac{5}{2} \right]
$$
$$
S_{\text{rot}} = k \left[ \ln \left( \frac{\sqrt{\pi}}{\sigma} \left( \frac{8\pi^2 kT}{h^2} \right)^{3/2} \sqrt{I_A I_B I_C} \right) + \frac{3}{2} \right]
$$
$$
S_{\text{vib}} = k \sum_i \left[ \frac{h\nu_i/kT}{\exp(h\nu_i/kT) - 1} - \ln \left( 1 - \exp(-h\nu_i/kT) \right) \right]
$$
From there we have entropy and enthalpy terms. We will now be able to calculate our Gibbs Free Energy over different temperatures using this equation
$$
G = H - TS
$$
## Interactive Results Viewer
Explore optimized molecular geometries with optimization convergence plots, vibrational frequencies, and thermochemical properties.
# RevGRN
Source: https://docs.revilico.bio/docs/revgrn
Gene Regulatory Network Inference and Boolean Attractor Simulation from Expression Data
## Why Use This Engine?
In the documentation below, we will use Revilico's RevGRN engine to infer gene regulatory networks from single-cell, bulk, or spatial transcriptomic data and simulate the stable cellular states that emerge from those regulatory interactions. RevGRN identifies which transcription factors are driving gene expression changes in the dataset, constructs a network of TF-target regulatory edges, and then uses Boolean dynamical simulation to identify the attractor states that the network converges to, enabling prediction of how genetic perturbations such as knockouts or overexpression events alter the cellular phenotype.
## Background
Gene regulatory networks describe the control logic by which transcription factors bind to gene promoters and enhancers to activate or repress downstream target genes. These networks determine which genes are expressed in each cell type and how expression patterns change in response to developmental signals, disease mutations, or drug perturbations. Inferring GRNs from transcriptomic data is challenging because correlation in expression does not imply causation, and regulatory relationships must be distinguished from indirect co-expression driven by shared upstream regulators.
RevGRN addresses this using machine learning-based importance scoring (GRNBoost2) to prioritize direct TF-target regulatory edges, filtering to the most informative subset of the network, and then converting the inferred network into a Boolean dynamical model to simulate cellular states. Boolean models represent each gene as either ON or OFF and apply the regulatory logic iteratively until the system converges to stable attractor states. These attractors correspond to distinct biological phenotypes (cell types, disease states, drug-response signatures), and perturbation simulations reveal how knocking out or overexpressing a gene redirects the network toward different attractors.
## Input Data Loading
RevGRN accepts expression matrices in CSV, TSV, H5AD, and LOOM formats. The engine auto-detects whether rows represent cells and columns represent genes, or vice versa, by analyzing the index patterns: Ensembl IDs (ENSG prefix), gene symbols, or cell barcode patterns are detected and the matrix is transposed if necessary. For matrices indexed by Ensembl IDs with a gene\_name column, symbols are resolved via the MyGeneInfo API.
## Quality Control and Normalization
Cells with fewer than `min_genes_per_cell` expressed genes (default 200) are removed. Genes detected in fewer than `min_cells_per_gene` cells (default 3) are removed. The filtered matrix is then normalized by library size and log-transformed:
$$
\tilde{x}_{ij} = \log\left(\frac{x_{ij}}{\sum_j x_{ij}} \cdot 10000 + 1\right)
$$
Where $x_{ij}$ is the raw count for gene $j$ in cell $i$. Highly variable genes are then selected by ranking genes on their dispersion ratio:
$$
d_j = \frac{\text{Var}(x_j)}{\text{mean}(x_j) + \varepsilon}
$$
The top 2,000 genes by dispersion (or adaptively capped at $\min(2000, 20 \times n_{\text{cells}})$) are retained for network inference.
## GRN Inference
**GRNBoost2 (Primary Method)**
GRNBoost2 uses an ensemble of gradient-boosted regression trees to estimate the regulatory importance of each transcription factor for each target gene. For each target gene, an XGBoost model is trained to predict that gene's expression from the expression of all TF genes. The feature importances from the trained model quantify how much each TF reduces prediction error, providing a directed importance score for each TF-target pair. GRNBoost2 is fast, parallelizable, and robust to non-linear regulatory relationships.
**Mutual Information (Alternative)**
Mutual information between TF and target gene expression is computed as:
$$
I(X; Y) = H(Y) - H(Y | X)
$$
Where $H(Y)$ is the marginal entropy of target gene expression and $H(Y|X)$ is the conditional entropy given TF expression. Mutual information captures non-linear associations that linear correlation misses.
**Spearman Correlation (Fallback)**
When neither GRNBoost2 nor mutual information libraries are available, Spearman rank correlation is used as a fallback. The absolute correlation value serves as the regulatory importance score.
After inference, the top `max_network_edges` edges (default 50) ranked by importance score are retained as the final network.
## Boolean Simulation
The inferred network is converted to a Boolean dynamical model. Each gene is represented as a binary variable (ON = 1, OFF = 0). At each simulation step, the state of each gene is updated synchronously according to the regulatory logic:
$$
s_j^{(t+1)} = \begin{cases} 1 & \text{if } \sum_i w_{ij} \cdot s_i^{(t)} > 0 \\ 0 & \text{otherwise} \end{cases}
$$
Where $w_{ij}$ is the regulatory edge weight from TF $i$ to gene $j$ (positive for activation, negative for repression, using the sign of the inferred importance). The simulation is initialized from random binary starting states and iterated until convergence (a state that maps to itself) or a limit of 50 steps. Each terminal state is recorded as an attractor. The simulation is repeated from 200 different random initial states, and the basin of attraction for each attractor is estimated as the fraction of initial states that converge to it.
## Perturbation Analysis
In knockouts, the target gene is pinned to state 0 (OFF) throughout the simulation regardless of its regulatory inputs. In overexpression simulations, the gene is pinned to state 1 (ON). The resulting attractor distribution is compared to the baseline to predict how the perturbation redirects network dynamics. Genes whose knockout or overexpression produces the largest shift in attractor basin fractions are the regulatory hubs most relevant to phenotypic control.
## Running the Engine
### Inputs
| Parameter | Default | Description |
| ---------------------------- | -------------------- | ---------------------------------------------------------------- |
| Expression matrix | Required | CSV, TSV, H5AD, or LOOM (cells x genes) |
| Data type | scrna | `scrna`, `bulk`, or `spatial` |
| Organism | human | `human`, `mouse`, or `rat` for TF annotation |
| Min genes per cell | 200 | QC threshold for sparse cell removal |
| Min cells per gene | 3 | QC threshold for sparse gene removal |
| Highly variable genes | 2000 | Number of HVGs for network inference |
| Max network edges | 50 | Top TF-target edges to retain |
| Inference method | grnboost2 | `grnboost2`, `mutual_info`, or `correlation` |
| Simulation type | attractor\_landscape | `attractor_landscape`, `gene_knockout`, or `gene_overexpression` |
| Knockout/overexpression gene | None | Gene symbol to perturb |
### Outputs
* **Network:** Ranked list of TF-target regulatory edges with importance scores
* **TF hub ranking:** Top transcription factors by total regulatory influence (sum of outgoing edge importance)
* **Attractors:** Stable cellular states with active gene sets and basin of attraction fractions
* **Perturbation comparison:** Baseline vs. perturbed attractor distribution (if perturbation mode selected)
* **Network visualization:** Interactive force-directed graph with TFs as blue diamonds and target genes as gray nodes, edge width proportional to importance
* **QC summary:** Cell counts, gene counts, HVG count, network edge count, TF count, and attractor count
* **Downloads:** Network CSV (TF, target, importance), attractors CSV (active genes and basin fraction), full results JSON
# RevMD-Aqua
Source: https://docs.revilico.bio/docs/revmd-aqua
Protein-Water Molecular Dynamics for Conformational Sampling and Binding Site Analysis
## Why Use This Engine?
In the documentation below, we will use Revilico's RevMD-Aqua engine to simulate a protein in explicit solvent under physiologically relevant conditions, without any ligand present. This type of simulation captures the natural conformational dynamics of the protein, revealing how the structure breathes, which binding site regions are flexible, and which protein conformations are thermodynamically accessible. The resulting trajectory snapshots serve as input for ensemble docking and as a reference for assessing target druggability under dynamic rather than static conditions.
## Background
Protein structures determined by crystallography or cryo-EM represent a single low-energy snapshot of a dynamic system. In solution, proteins undergo continuous conformational fluctuations ranging from local side chain rotations to domain-level rearrangements, and these motions are fundamentally important for function, allostery, and ligand binding. A binding site that appears druggable in a crystal structure may be transiently open or closed in solution, and a cryptic site invisible in the static structure may be revealed during molecular dynamics. RevMD-Aqua simulates the protein in a periodic water box using GROMACS, generating a trajectory that captures this conformational diversity over nanosecond timescales.
## Simulation Pipeline
**System Construction**
The protein structure is placed in a periodic simulation box, typically a dodecahedron or cubic geometry, with a minimum distance of 1.0 nm from the protein surface to the box edge in any direction. The box is then filled with explicit TIP3P water molecules and neutralized by adding sodium and chloride ions to a physiological concentration of 0.15 M NaCl. This solvation environment mimics the ionic strength and dielectric properties of the cellular aqueous environment.
**Force Field Assignment**
Each atom in the system is assigned parameters from the AMBER99SB-ILDN force field, which defines the potential energy function governing all interatomic interactions. The total potential energy of the system is the sum of bonded and non-bonded terms:
$$
E_{\text{total}} = E_{\text{bonds}} + E_{\text{angles}} + E_{\text{dihedrals}} + E_{\text{VDW}} + E_{\text{elec}}
$$
The bonded terms use harmonic potentials for bonds ($k_b(r - r_0)^2$) and angles ($k_\theta(\theta - \theta_0)^2$) and a periodic potential for dihedrals. Non-bonded interactions use a Lennard-Jones potential for van der Waals forces and a Coulomb potential for electrostatics, both computed with particle-mesh Ewald (PME) summation for long-range accuracy. Forces on each atom are computed as $F_i = -\nabla_{r_i} E$, and Newton's equations of motion are integrated using the leapfrog algorithm with a 2 fs timestep.
**Energy Minimization**
Before dynamics begin, a steepest descent energy minimization step removes steric clashes and unrealistic geometries introduced during system construction. The system is minimized until the maximum force on any atom falls below 1000 kJ/mol/nm.
**Equilibration**
Equilibration occurs in two phases. In the NVT phase (constant volume and temperature), the system is heated to 300 K using a velocity-rescaling thermostat while restraining protein heavy atoms to their initial positions. This allows water molecules and ions to relax around the protein surface. In the NPT phase (constant pressure and temperature), the pressure barostat (Parrinello-Rahman) is activated and the box dimensions equilibrate to the correct density at 1 bar. Both phases run for 100 ps.
**Production MD**
The unrestrained production simulation runs for the user-specified duration (default 100 ns). Coordinates are written to the trajectory file at regular intervals (typically every 10 ps), producing a set of protein snapshots that capture the natural conformational ensemble. Velocities are initialized from a Maxwell-Boltzmann distribution at 300 K.
## Analysis Outputs
**RMSD (Root Mean Square Deviation)**
RMSD measures how much the protein structure deviates from the initial reference frame over time. It is computed for all backbone atoms after least-squares superposition:
$$
\text{RMSD}(t) = \sqrt{\frac{1}{N}\sum_{i=1}^{N} |r_i(t) - r_i^{\text{ref}}|^2}
$$
A stable RMSD plateau indicates that the simulation has converged to an equilibrium ensemble. Rising RMSD indicates ongoing structural rearrangement.
**RMSF (Root Mean Square Fluctuation)**
RMSF measures the time-averaged positional flexibility of each residue around its mean position. High RMSF regions correspond to flexible loops, termini, and dynamic binding site elements. Low RMSF regions correspond to the rigid hydrophobic core.
**Radius of Gyration**
The radius of gyration $R_g = \sqrt{\sum m_i |r_i - r_{\text{cm}}|^2 / \sum m_i}$ monitors overall protein compactness over the simulation. Systematic changes in $R_g$ can indicate unfolding or large-scale conformational transitions.
**Binding Site Dynamics**
For each identified binding pocket, volume and accessibility are tracked over the trajectory using the same Voronoi tessellation approach applied in RevPocket. This reveals whether pockets are constitutively open, transiently formed, or progressively closed, directly informing the value of ensemble docking against the extracted snapshots.
**Trajectory Snapshots for Ensemble Docking**
The production trajectory is sampled at regular intervals to extract protein conformations for use as receptor inputs in RevMD-Bind ensemble docking. Snapshots are structurally clustered and representative frames from each cluster are exported as PDB files.
## Running the Engine
### Inputs
| Parameter | Default | Description |
| ----------------- | -------------- | -------------------------------------------------- |
| Protein PDB | Required | Input protein structure (no ligand) |
| Force field | amber99sb-ildn | Molecular mechanics force field |
| Water model | tip3p | Explicit solvent model |
| Box type | dodecahedron | Simulation box geometry |
| Box distance | 1.0 nm | Minimum protein-to-edge distance |
| Ion concentration | 0.15 M | NaCl concentration |
| Simulation length | 100 ns | Total production MD duration |
| Output interval | 10 ps | Trajectory write frequency |
| Snapshot interval | 10 ns | Frequency for ensemble docking snapshot extraction |
### Outputs
* **Trajectory file:** Full atomic trajectory for visualization and analysis
* **RMSD plot:** Backbone deviation from the starting structure over time
* **RMSF plot:** Per-residue flexibility colored on the 3D structure
* **Radius of gyration plot:** Protein compactness over simulation time
* **Binding site volume plot:** Pocket opening and closing dynamics
* **Ensemble snapshots:** Clustered PDB frames for downstream ensemble docking input
* **Energy components:** Potential energy, temperature, pressure, and density over time
# RevMD-Bind
Source: https://docs.revilico.bio/docs/revmd-bind
Understanding Drug-Target Dynamics with Revilico's MD Simulation Suite: Protein, Ligand, and Membrane Systems
## Why Use this product?
Molecular Dynamics (MD) simulations reveal how biomolecular systems evolve over time in their native environments, capturing the dynamic behavior that static structures miss, from protein flexibility and binding site breathing (protein-water), to ligand binding stability and key interaction identification (protein-ligand), to membrane permeability and drug-lipid interactions (ligand-membrane). These simulations bridge the gap between static docking predictions and experimental validation, helping drug discovery teams prioritize compounds by assessing binding pose stability, understanding target flexibility, and predicting ADME properties like membrane permeability. By providing atomic-level insight into molecular motion and interactions over nanosecond-to-microsecond timescales, MD enables structure-based drug design to move from static snapshots into dynamic, physically realistic predictions.

## Background
Proteins, ligands and membranes are dynamic systems that continuously fluctuate, reorganize, and respond to their environment. Where tools like Docking cannot capture this system and can only capture static snapshots of these interactions, Molecular Dynamics (MD) proves to be a solution by providing a physics based framework for understanding how biomolecular systems behave over time. Now what does Molecular Dynamics actually do? At its core, MD simulates how atoms move as a function of time by applying classical mechanics. Generally MD follows this corresponding pipeline. First system construction places (proteins, ligands, and/or lipids) in a periodic simulation box, typically solvated with explicit water molecules and ions to mimic physiological conditions. Second we assign a force field. These are the parameters assigned to each atom (i.e. partial charge, bonded connectivity, nonbonded interaction parameters, etc.). Our third step is energy minimization. This step removes steric clashes or unrealistic geometries that could cause numerical instability. Our next step is bringing the system to equilibration, both to a target temperature and allowing the system density and box dimension equilibrate to realistic pressure. The next step is production molecular dynamics where the system is simulated over time, generating a trajectory that captures atomic motion under realistic conditions.
Now we can dive into the theory. The behavior of our system is governed by a classical potential energy function or the force file which is expressed with the following equation.
$$
E = \sum \text{interactions}
$$
$$
E_{\text{total}} = E_{\text{bonds}} + E_{\text{angle}} + E_{\text{dihedral}} + E_{\text{VDW}} + E_{\text{elec}}
$$
$$
E_{\text{bond}} = \sum k_b (r - r_0)^2
$$
$$
E_{\text{angle}} = \sum k_0 (\theta - \theta_0)^2
$$
$$
E_{\text{dihedral}} = \sum \frac{V_n}{2} (1 + \cos(n\phi - \gamma))
$$
$$
E_{\text{VDW}} = \sum 4\epsilon \left[ \left( \frac{\sigma}{r} \right)^{12} - \left( \frac{\sigma}{r} \right)^6 \right]
$$
$$
E_{\text{elec}} = \sum (q_i \times q_j)/(4\pi \epsilon_0 r_{ij})
$$
From the $E_{\text{total}}$ function, forces acting on each atom are computed as
$$
F_i = - \nabla_{r_i} E
$$
Using Newton’s equation F = ma, we are then able to use this calculated force and the mass of each atom to update positions over time by updating our velocity that results from acceleration from this equation.
With temperature in MD, it represents the average kinetic energy of atoms. Initial velocities are drawn from a Maxwell-Boltzmann distribution, and thermostats are used to maintain the desired temperature statistically over time. .Thermal motion ensures that atoms continuously explore conformational space, allowing the system to sample biologically relevant states rather than freezing into a single configuration
**Running Snapshot/Post-Processing Free Energy Methods**
For MMPBSA, we are evaluating whether is it preferred for the ligand to be docked in the protein or if it is energetically favorable for the ligand to be free. In order to calculate MMPBSA we do the following: 1) we take note that we have our entire MD simulation made up of all the snapshots based on our intervals across the total simulation time.. We will remove the early snapshots as this is the equilibration step and therefore are not representative samples for calculating energies, and ensure that our calculation does not include extreme values that can heavily skew our outcomes. From there, we take the snapshots at the defined intervals and we calculate the following equation at each time step and average out all the values across each snapshot/interval.
$$
\Delta G_{\text{bind}} = G_{\text{complex}} - (G_{\text{protein}} + G_{\text{ligand}})
$$
$$
G \approx E_{\text{MM}} + G_{\text{solv}}
$$
Where $E_{\text{MM}}$ is the forcefield equation defined above as $E_{\text{total}}$ Or $\Delta G_{\text{bind}}$. When we decompose this equation, essentially we will need to compute the energetics using our selected force fields for the entire complex (protein with ligand docked), the energies of the force fields of the protein on its own (internal forces of the protein atoms acting upon itself), and the force field of the ligand on its own (forces of the ligand atoms acting upon itself). Furthermore we can decompose $G_{\text{solv}}$ to the following equation:
$$
G_{\text{solv}} = G_{\text{polar}}^{\text{PB}} + G_{\text{nonpolar}}^{\text{SA}}
$$
Where $G_{\text{polar}}^{\text{PB}}$ is the electrostatic cost of putting a charged molecule in water and $G_{\text{nonpolar}}^{\text{SA}}$ is the hydrophobic cost of creating a cavity in water. Essentially what we are calculating is the free energy cost of whether the ligand is better in the bound position or whether it is better unbound. How we can interpret this is that very low delta G\_total values indicate strong preference for the bound state which is the preferred value when it comes to evaluating whether a compound is a strong candidate. These energetic calculations elucidate a more sophisticated binding composition, breaking down all of the energetic contributions to direct ligand protein engagements.
## Interactive Results Viewer
Explore Protein-Ligand MD simulation results interactively. View trajectory analysis, RMSD plots, energy decomposition, and MMGBSA binding free energy contributions.
# RevMD-Mem
Source: https://docs.revilico.bio/docs/revmd-mem
Ligand-Membrane Molecular Dynamics for Permeability and Drug-Lipid Interaction Analysis
## Why Use This Engine?
In the documentation below, we will use Revilico's RevMD-Mem engine to simulate a small molecule compound interacting with a lipid bilayer membrane, providing insight into passive membrane permeability, lipid partitioning, and drug-lipid interactions. These properties are directly relevant to ADME prediction: a compound that cannot permeate cellular membranes cannot reach intracellular targets regardless of binding affinity. RevMD-Mem quantifies the free energy cost of membrane crossing, enabling the prioritization of compounds with favorable permeability profiles before experimental testing.
## Background
The lipid bilayer is a selectively permeable barrier composed of two leaflets of phospholipid molecules arranged with their hydrophilic headgroups facing aqueous solution and their hydrophobic acyl tails forming the membrane interior. Passive membrane permeation requires a compound to partition from the aqueous phase into the hydrophobic core, traverse the bilayer, and re-partition into the aqueous phase on the opposite side. The free energy profile along this permeation pathway, known as the potential of mean force (PMF), directly determines the permeation rate and membrane residence time.
Molecular dynamics simulation of a compound within an explicit lipid bilayer provides an atomistically detailed view of drug-membrane interactions that continuum solubility models cannot capture. RevMD-Mem uses the GROMACS simulation engine with the CHARMM36 force field for lipids and the CGenFF (CHARMM General Force Field) for small molecules, generating PMF profiles and partition coefficients that predict passive permeability.
## Simulation Pipeline
**Membrane and System Construction**
A pre-equilibrated lipid bilayer patch (POPC by default, or DPPC for saturated membrane models) is used as the starting membrane geometry. The bilayer is oriented with the membrane normal along the z-axis. The compound of interest is initially placed in the aqueous phase at a specified distance from the bilayer surface. A periodic water box is added above and below both leaflets, and ions are added to 0.15 M NaCl. The resulting system contains approximately 128 to 256 lipid molecules, approximately 8,000 to 16,000 water molecules, and one or more drug molecules.
**Force Field Parameters**
Lipid parameters are taken from the CHARMM36 lipid force field, which has been extensively validated against experimental structural and thermodynamic data for biological membranes. Small molecule parameters are generated using the CGenFF parameterization tool, which assigns atom types and partial charges by analogy to known fragments in the CGenFF library. The same potential energy function as RevMD-Bind governs all interactions:
$$
E_{\text{total}} = E_{\text{bonds}} + E_{\text{angles}} + E_{\text{dihedrals}} + E_{\text{VDW}} + E_{\text{elec}}
$$
Long-range electrostatics are handled with particle-mesh Ewald (PME) summation. A semi-isotropic pressure coupling (Parrinello-Rahman) is applied so that the membrane area and thickness can relax independently, which is essential for correct bilayer mechanics.
**Equilibration**
System equilibration follows the same NVT then NPT protocol as RevMD-Aqua, with heavy atoms initially restrained. The membrane-specific NPT equilibration uses semi-isotropic pressure coupling to allow the bilayer area per lipid to relax to its equilibrium value while maintaining the correct surface tension.
**Umbrella Sampling for PMF Calculation**
To compute the free energy profile along the membrane permeation pathway, the compound is driven through the bilayer using umbrella sampling. A series of simulation windows is generated by placing the compound at evenly spaced positions along the z-axis from the aqueous phase through the membrane center and out the other side. The spacing between windows is 0.1 to 0.2 nm, producing 30 to 50 windows covering the full bilayer thickness. In each window a harmonic restraining potential:
$$
V_{\text{umbrella}}(z) = \frac{1}{2} k_{\text{pull}} (z - z_0)^2
$$
keeps the compound near its target z-position $z_0$ with a spring constant $k_{\text{pull}}$ of approximately 1000 kJ/mol/nm$^2$. Each window is simulated for 5 to 10 ns to sample the local free energy landscape.
**WHAM (Weighted Histogram Analysis Method)**
The biased distributions from all umbrella windows are combined using WHAM to reconstruct the unbiased potential of mean force:
$$
F(z) = -k_BT \ln \rho(z)
$$
Where $\rho(z)$ is the unbiased probability density of the compound at position z along the bilayer normal. The PMF is computed iteratively until the free energy estimates converge. The resulting profile shows the free energy barriers to membrane entry, the partition free energy into the hydrophobic core (log $P_{\text{mem/water}}$), and the overall free energy cost of traversing the bilayer.
## Analysis Outputs
**Potential of Mean Force Profile**
The PMF plot shows the free energy in kcal/mol as a function of position along the membrane normal. Key features include the barrier height at the membrane-water interface (related to initial partitioning), the depth of the free energy minimum in the hydrophobic core (the membrane affinity), and the overall permeation barrier from one aqueous phase to the other. A low, broad minimum with small interfacial barriers characterizes lipophilic compounds that partition efficiently into membranes. A high central barrier with a deep minimum characterizes "membrane-trapped" compounds that partition into but do not easily permeate the bilayer.
**Log Kp (Membrane-Water Partition Coefficient)**
The free energy minimum relative to the aqueous bulk gives the membrane-water partition coefficient:
$$
\log K_p = \frac{-\Delta G_{\text{min}}}{k_BT \ln 10}
$$
**Permeability Coefficient**
The overall passive permeability is estimated from the PMF and diffusion coefficient profiles using the inhomogeneous solubility-diffusion model:
$$
P = \left[ \int_{-d/2}^{d/2} \frac{e^{F(z)/k_BT}}{D(z)} \, dz \right]^{-1}
$$
Where $D(z)$ is the position-dependent diffusion coefficient estimated from the autocorrelation of coordinate fluctuations in each umbrella window.
**Lipid Interaction Analysis**
The compound's preferential interactions with specific lipid headgroup atoms and acyl chain carbons are quantified through radial distribution functions and contact analysis, identifying whether the compound intercalates into the hydrophobic core or associates preferentially with specific lipid regions.
## Running the Engine
### Inputs
| Parameter | Default | Description |
| ---------------------------- | --------------- | ------------------------------------------ |
| Compound SMILES or SDF | Required | Small molecule structure |
| Membrane type | POPC | Lipid composition (POPC, DPPC, or mixed) |
| Force field | CHARMM36 | Lipid force field |
| Small molecule FF | CGenFF | Ligand parameterization method |
| Simulation length per window | 5 ns | Production time per umbrella window |
| Window spacing | 0.1 nm | Distance between umbrella sampling windows |
| Pull force constant | 1000 kJ/mol/nm2 | Harmonic restraint spring constant |
| Temperature | 310 K | Physiological temperature |
### Outputs
* **PMF profile:** Free energy along membrane normal with statistical uncertainty
* **Log Kp:** Membrane-water partition coefficient
* **Permeability coefficient:** Estimated passive permeability in cm/s
* **Bilayer thickness and area per lipid:** Structural properties during simulation
* **Compound tilt and orientation plots:** Preferred molecular orientation within the bilayer
* **Lipid contact analysis:** Preferential interactions with specific lipid components
* **Convergence analysis:** PMF bootstrap uncertainty and per-window sampling quality
# RevMethyl
Source: https://docs.revilico.bio/docs/revmethyl
Epigenetic Biological Age Estimation from DNA Methylation Using Validated Epigenetic Clocks
## Why Use This Engine?
In the documentation below, we will use Revilico's RevMethyl engine to compute epigenetic biological age from DNA methylation beta values using four validated published clocks. This engine enables researchers and clinicians to measure how fast a sample is aging biologically, independent of its chronological age, and to quantify the degree of age acceleration or deceleration relative to population norms.
## Background
DNA methylation (DNAm) is a heritable epigenetic modification in which a methyl group is added to the cytosine of a CpG dinucleotide. Methylation levels are quantified as beta values, representing the proportion of cells in a sample in which a given CpG site is methylated, ranging continuously from 0 (fully unmethylated) to 1 (fully methylated). DNAm patterns change predictably with age across a large fraction of the genome, and this property was exploited to build regression models, known as epigenetic clocks, that predict chronological or biological age from a weighted linear combination of CpG site beta values.
Epigenetic clocks are among the most accurate and reproducible biomarkers of biological aging available. They capture aspects of aging that are not reflected by chronological age alone, including accumulated cellular damage, lifestyle exposures, disease burden, and mortality risk. RevMethyl implements four widely used and validated clocks: the Horvath 2013 pan-tissue clock, the Hannum 2013 blood clock, the PhenoAge 2018 phenotypic age clock, and the GrimAge 2019 mortality clock. Each clock was trained on different biological outcomes and covers a distinct number of CpG probes, and their results together provide a multi-dimensional picture of epigenetic aging.
## Input and Data Loading
RevMethyl accepts a two-column CSV or TSV file where the first column contains Illumina 450k or EPIC array probe identifiers (e.g. `cg16867657`) and the second column contains the corresponding beta value for each probe. Gzip-compressed files (`.csv.gz`, `.tsv.gz`) are also accepted. A header row is auto-detected and skipped if the first cell does not begin with a "cg" prefix. Beta values are clipped to the valid range \[0, 1] prior to computation, and any out-of-range values are flagged in the QC statistics.
Alternatively, raw IDAT files (Red and Green channel) can be uploaded directly. The engine processes the IDAT pair through array annotation to derive probe-level beta values before passing them to the clock computation step. The array type (450k, EPIC, or EPICv2) can be specified or detected automatically.
## Quality Control
Before computing any clock, the engine collects the following QC statistics from the input beta matrix: total number of probes present, mean beta value across all probes, standard deviation of beta values, and minimum and maximum observed beta values. These statistics serve as a data sanity check. For each clock, the engine also reports the number of probes matched from the clock definition, the number of probes that were absent in the input file and required imputation, and the resulting probe coverage fraction.
## Epigenetic Clock Computation
All four clocks share the same underlying computation structure. Each clock defines a set of CpG probes with associated regression coefficients and a scalar intercept. The raw clock score is computed as:
$$
\text{score} = \beta_0 + \sum_{i=1}^{n} w_i \cdot x_i
$$
Where $\beta_0$ is the clock-specific intercept, $w_i$ is the regression coefficient for probe $i$, and $x_i$ is the observed beta value for probe $i$. After this linear combination, each clock either applies a transformation or returns the raw score directly as DNAm age, depending on the clock's original training procedure. The final age estimate is clipped to the biologically plausible range of \[0, 120] years.
**Missing Probe Imputation**
Not all input datasets will contain every probe defined by a given clock. When a clock probe is absent from the input file, the engine imputes its beta value using the mean beta of all probes that were successfully matched for that clock. This is the standard approach used in the original clock publications and ensures that the score degrades gracefully with decreasing array coverage rather than failing entirely. Each missing probe contributes $w_i \cdot \bar{x}_{\text{matched}}$ to the raw score, where $\bar{x}_{\text{matched}}$ is the mean of matched probe beta values.
**Horvath 2013**
The Horvath clock is a pan-tissue clock trained on 51 different tissue and cell types using penalized regression (elastic net) with 353 CpG probes. It is the most broadly applicable of the four clocks and was designed to predict age consistently regardless of tissue origin. The clock uses an intercept of 0.6955. Because the clock was trained with a non-linear age transformation applied to the outcome variable (compressing the age scale in childhood), the raw score must be inverted through a piecewise anti-transformation before reporting DNAm age:
$$
\text{DNAm age} = \begin{cases} 21 \cdot e^{\text{score}} - 1 & \text{if score} < 0 \\ 21 \cdot (\text{score} + 1) - 1 & \text{if score} \geq 0 \end{cases}
$$
The negative branch uses an exponential to expand compressed childhood ages, while the non-negative branch uses a linear mapping for adult ages. The result is a DNAm age estimate in years.
**Hannum 2013**
The Hannum clock is a blood-specific linear clock trained on whole-blood methylation from 71 CpG probes using ridge regression with chronological age as the outcome. It uses an intercept of 0.0 and returns the raw score directly as DNAm age without any transformation. Due to its small probe set and tissue specificity, it is the fastest clock to compute and is most accurate when the input sample is blood-derived.
**PhenoAge 2018**
PhenoAge uses 513 CpG probes and was trained not on chronological age directly but on a composite phenotypic age score derived from clinical biomarkers including albumin, creatinine, glucose, C-reactive protein, lymphocyte percentage, mean cell volume, red cell distribution width, alkaline phosphatase, and white blood cell count, combined with mortality risk modeling. The result is a DNAm age estimate that reflects morbidity and mortality risk more strongly than calendar age. PhenoAge uses an intercept of 60.664 and returns the raw score directly. A higher PhenoAge relative to chronological age indicates elevated biological aging associated with chronic disease burden.
**GrimAge 2019**
GrimAge is the most complex of the four clocks, using 1,030 CpG probes. Rather than being trained directly on age or phenotypic age, GrimAge is a composite of DNAm-based surrogate scores for several plasma proteins and smoking pack-years. These surrogates were each individually trained using elastic net regression, and their weighted combination was then optimized to predict time-to-death from all causes. GrimAge uses an intercept of 25.0 and returns the raw composite score directly as DNAm age. Among the four clocks, GrimAge is the strongest predictor of mortality and disease onset, making it the most clinically informative metric for assessing biological age acceleration.
## Age Acceleration
If the user provides the sample's chronological age, the engine computes age acceleration for each clock:
$$
\text{Age Acceleration} = \text{DNAm Age} - \text{Chronological Age}
$$
Positive age acceleration indicates that the sample is epigenetically older than its calendar age, reflecting faster biological aging. Negative age acceleration indicates slower biological aging. Age acceleration is the primary metric used to identify samples with atypical aging trajectories relative to population norms.
## Running the Engine
### Inputs
| Parameter | Required | Description |
| ------------------- | -------- | ----------------------------------------------------------- |
| `methylation_file` | Yes | Two-column CSV or TSV with probe IDs and beta values |
| `chronological_age` | No | Decimal age of the sample for acceleration calculation |
| `clocks` | No | Comma-separated subset of clocks to run (default: all four) |
### Outputs
Upon completion, the engine returns the following for each clock:
* **DNAm age:** Estimated biological age in years
* **Age acceleration:** DNAm age minus chronological age (if chronological age provided)
* **Probe coverage:** Fraction of clock-defined probes matched in the input file
* **Matched probes:** Count of probes found in the input
* **Imputed probes:** Count of probes absent from the input and imputed with mean beta
Summary statistics across all clocks are also returned: mean DNAm age, mean age acceleration, mean probe coverage, and number of clocks computed.
**QC metrics** include total probes in the input file, mean and standard deviation of beta values, and minimum and maximum beta values observed.
# Migration / Invasion
Source: https://docs.revilico.bio/docs/revmigration-invasion
Cancer Cell Motility and Invasive Capacity Prediction Under Drug Treatment
## Why Use This Engine?
In the documentation below, we will use Revilico's Migration and Invasion assay engine to predict the effect of drug treatment on cancer cell motility and invasive capacity across cell lines with different inherent metastatic potentials. This assay is directly relevant to anti-metastatic drug development: a compound that kills cells in primary culture may or may not inhibit the migratory behavior that drives metastatic dissemination, and these two properties require independent evaluation.
## Background
Cancer metastasis requires tumor cells to detach from the primary tumor, migrate through the extracellular matrix, invade through basement membranes, and establish colonies at distant sites. Cell migration is measured in vitro using two complementary formats: the scratch (wound healing) assay measures how quickly cells close a mechanically created gap in a cell monolayer, and the Transwell invasion assay measures how many cells actively migrate through a porous membrane coated with Matrigel (a reconstituted basement membrane extract) toward a chemoattractant.
Cell lines differ substantially in their baseline migratory capacity based on their epithelial-mesenchymal transition (EMT) status. MDA-MB-231 (triple-negative breast cancer) is highly mesenchymal with a high basal migration rate. MCF7 (luminal breast cancer) retains an epithelial phenotype and migrates slowly. RevAssay models both assay formats using cell-line-specific basal migration rates and the GNN+MLP viability signal as the drug-induced suppression factor.
## Simulation Model
**Basal migration rates** by cell line (calibrated to experimental scratch assay literature values):
| Cell Line | Basal rate | Phenotype |
| ---------- | ---------- | ----------------------------------------- |
| MDA-MB-231 | 22 um/h | Highly mesenchymal, metastatic TNBC |
| A549 | 14 um/h | Moderately invasive lung adenocarcinoma |
| MCF7 | 9 um/h | Low-invasiveness epithelial breast cancer |
| HEPG2 | 6 um/h | Low motility hepatocellular carcinoma |
The drug-suppressed migration rate scales proportionally with predicted viability:
$$
\text{rate}(c) = R_0 \cdot v(c) \cdot \eta, \quad \eta \sim \mathcal{LN}(0, 0.15^2)
$$
Where $R_0$ is the cell-line basal rate and $v(c)$ is the GNN+MLP viability at concentration $c$.
**Scratch assay wound closure:**
$$
\text{Closure\%}(c) = v(c) \times 85 + \varepsilon, \quad \varepsilon \sim \mathcal{N}(0, 8^2)
$$
**Transwell invasion index** (Matrigel-coated membrane):
$$
\text{Invasion index}(c) = v(c) \times 0.78 \cdot \eta_{\text{inv}}, \quad \eta_{\text{inv}} \sim \mathcal{LN}(0, 0.12^2)
$$
The invasion index is lower than wound closure because Matrigel provides an additional physical barrier that requires active matrix metalloproteinase (MMP) activity to penetrate.
**Kinetic time series** for scratch assay wound closure over the time course $T$:
$$
\text{Closure\%}(t) = v(c) \times 90 \times \left(1 - e^{-3t/T}\right) + \varepsilon(t)
$$
This exponential approach to the asymptote reflects the decelerating rate of gap closure as the wound narrows.
## Parameters
| Parameter | Default | Description |
| ---------------------- | ------- | ------------------------------------------------ |
| Assay type | Scratch | `Scratch` (wound healing) or `Transwell` |
| Matrigel | Off | Enable Matrigel coating for invasion measurement |
| Time course | 72 h | Duration of kinetic imaging (24 to 120 hours) |
| Exposure hours | 72 h | Drug exposure duration |
| Hill coefficient | 1.2 | Dose-response steepness |
| Biological variability | None | Noise level for replicates |
## Outputs
* **Migration rate (um/h):** Drug-suppressed migration rate per cell line per concentration
* **Wound closure %:** Percentage of scratch gap closed at the assay endpoint
* **Invasion index:** Normalized invasive capacity (Transwell mode only)
* **Kinetic curves:** Time-series wound closure from 0 to the full time course
* **Dose-response curves:** Migration rate and wound closure across the concentration range
* **Comparative bar chart:** Cell lines ranked by migration rate, colored by drug concentration
* **96-well heatmap:** Plate-view colored by migration rate or wound closure fraction
# RevMut-PMX
Source: https://docs.revilico.bio/docs/revmut-pmx
Alchemical Free Energy Calculations for Protein Mutation Binding Affinity Prediction
## Why Use This Engine?
In the documentation below, we will use Revilico's RevMut-PMX engine to predict how a single point mutation in a protein changes its binding affinity for a ligand. This type of calculation is directly applicable to resistance mutation modeling, where a clinical mutation in a drug target reduces binding of an approved inhibitor, and to protein engineering, where beneficial mutations are identified to improve binding of a therapeutic molecule. RevMut-PMX computes the relative binding free energy change (DDG\_bind) using PMX hybrid topology generation and GROMACS alchemical free energy perturbation simulations.
## Background
A point mutation changes one amino acid in the protein sequence, which alters the local electrostatic and steric environment of the binding site and thereby affects the binding free energy of any ligand in that site. Directly measuring this effect computationally would require running separate MD simulations of the wild-type and mutant proteins, both free and bound to the ligand, and computing absolute binding free energies for each. This is prohibitively expensive and converges poorly.
The thermodynamic cycle approach circumvents this by recognizing that binding free energy differences are path-independent. Rather than simulating the actual mutation event, alchemical free energy perturbation (FEP) simulates the unphysical but thermodynamically valid process of continuously transforming one amino acid residue into another within a single simulation, computing the free energy cost of this transformation both in the presence and absence of the ligand.
## Thermodynamic Cycle
The four corners of the thermodynamic cycle relate wild-type and mutant binding free energies:
$$
\Delta\Delta G_{\text{bind}} = \Delta G_{\text{mut}}^{\text{complex}} - \Delta G_{\text{mut}}^{\text{apo}}
$$
Where $\Delta G_{\text{mut}}^{\text{complex}}$ is the free energy of alchemically mutating the residue with the ligand present, and $\Delta G_{\text{mut}}^{\text{apo}}$ is the same mutation free energy in the protein alone without ligand. Because the thermodynamic cycle is closed, DDG\_bind can be obtained from these two alchemical legs without ever computing the physical binding event directly.
A negative DDG\_bind indicates that the mutation improves ligand binding (stabilizing). A positive DDG\_bind indicates that the mutation weakens binding (destabilizing), which is the signature of a resistance mutation.
## Hybrid Topology Generation (PMX)
RevMut-PMX uses PMX to generate a single hybrid topology that simultaneously represents both the wild-type residue (state A, lambda = 0) and the mutant residue (state B, lambda = 1). Atoms shared between the two residues retain their parameters throughout the simulation. Atoms present only in the wild type are converted to dummy atoms that gradually lose their physical interactions as lambda increases from 0 to 1. Atoms present only in the mutant are introduced as dummy atoms at lambda = 0 and gradually acquire their full interactions as lambda approaches 1. A soft-core potential prevents singularities in the van der Waals interactions when atoms appear or disappear:
$$
V_{\text{sc}}(r) = 4\epsilon \left[ \frac{1}{(\alpha \lambda^p \sigma^6 + r^6)^2} - \frac{1}{\alpha \lambda^p \sigma^6 + r^6} \right]
$$
Where alpha = 0.5, p = 1, and sigma = 0.3 nm are the soft-core parameters. Three lambda schedules are available: a standard 11-window schedule for typical mutations, a 17-window charge-changing schedule for mutations that alter the net charge (e.g., Asp to Asn), and a 21-window large-mutation schedule for drastic sidechain changes.
## Simulation Protocol
Each alchemical leg (complex and apo) follows the same simulation pipeline. After building the hybrid topology, the system is solvated in a dodecahedral water box with TIP3P water and 0.15 M NaCl. Energy minimization removes clashes, followed by 100 ps NVT and 100 ps NPT equilibration as described in the RevMD engines. Production FEP simulations run for the user-specified duration (default 5 ns) per lambda window. Both legs and all lambda windows can run in parallel.
During each production window, GROMACS computes the derivative of the Hamiltonian with respect to lambda (dH/dlambda) and the reduced potential at all other lambda states at every step, writing these to dhdl.xvg files that are read by the analysis step.
## Free Energy Analysis
After all lambda windows complete, the free energy difference for each leg is estimated by cascading through three estimators in order of accuracy:
**MBAR (Multistate Bennett Acceptance Ratio)** is the primary estimator. It uses the full matrix of reduced potentials across all lambda states simultaneously, extracting the maximum likelihood free energy differences:
$$
\Delta G = -k_BT \ln \frac{\sum_j \sum_n \frac{e^{-u_{kn}}}{\sum_l N_l e^{f_l - u_{ln}}}}{\sum_j \sum_n \frac{e^{-u_{jn}}}{\sum_l N_l e^{f_l - u_{ln}}}}
$$
**BAR (Bennett Acceptance Ratio)** is used if MBAR does not converge. It computes free energy differences between adjacent lambda windows using the acceptance probability of work values.
**TI (Thermodynamic Integration)** is used as a last resort. It integrates the mean dH/dlambda values across lambda using a trapezoid rule:
$$
\Delta G = \int_0^1 \left\langle \frac{\partial H}{\partial \lambda} \right\rangle_\lambda \, d\lambda
$$
The final DDG\_bind is computed as the difference between the complex and apo leg free energies, with uncertainty propagated as:
$$
\sigma_{\Delta\Delta G} = \sqrt{\sigma_{\text{complex}}^2 + \sigma_{\text{apo}}^2}
$$
## Running the Engine
### Inputs
| Parameter | Default | Description |
| -------------------------------- | ----------------------- | -------------------------------------------------------------------- |
| Protein PDB | Required | Wild-type protein structure |
| Mutation | Required | Format: `[WT_AA][RESNUM][MUT_AA]`, e.g. `N265M` |
| Chain ID | A | Protein chain containing the mutation |
| Ligand SDF/MOL2 | Optional | Ligand for complex leg (omit for stability-only calculation) |
| Pre-parameterized ligand GRO+ITP | Optional | Expert mode: supply GROMACS-ready ligand topology |
| Force field | amber99sb-star-ildn-mut | PMX-compatible force field |
| Water model | tip3p | Explicit solvent model |
| Lambda schedule | simple | `simple` (11 windows), `charge-changing` (17), `large-mutation` (21) |
| Simulation length | 5 ns | Production time per lambda window |
| Ion concentration | 0.15 M | NaCl concentration |
| Box type | dodecahedron | Periodic box geometry |
### Outputs
* **DDG\_bind:** Relative binding free energy change in kcal/mol with uncertainty estimate
* **DDG\_complex:** Free energy cost of mutation in the protein-ligand complex
* **DDG\_apo:** Free energy cost of mutation in the apo protein
* **Phase timeline:** Progress through prepare, setup, running, and analysis phases with timing
* **Convergence plots:** dH/dlambda or MBAR overlap matrices per leg
* **Interpretation:** Sign-annotated result with binding stabilizing or destabilizing classification
**DDG\_bind interpretation:**
| DDG\_bind | Classification |
| --------------------- | -------------------------------- |
| Below -1.5 kcal/mol | Strong binding improvement |
| -1.5 to -0.5 kcal/mol | Moderate improvement |
| -0.5 to +0.5 kcal/mol | Negligible effect |
| +0.5 to +1.5 kcal/mol | Moderate destabilization |
| Above +1.5 kcal/mol | Strong binding loss (resistance) |
# RevNotes
Source: https://docs.revilico.bio/docs/revnotes
Structured Research Note-Taking, Knowledge Capture, and Shared Scientific Documentation
## Why Use RevNotes?
RevNotes is Revilico's structured note-taking workspace for capturing scientific observations, hypotheses, experimental context, and project knowledge within the platform. Rather than maintaining separate lab notebooks or scattered documents, research teams use RevNotes to record findings alongside the computational data that generated them, share notes across the team, and publish structured knowledge documents for broader access.
## Interface
**Left navigation:**
* **All Notes:** Displays all notes created by the current user across all folders and projects
* **Shared Notes:** Notes shared with the current user by team members
**Top toolbar:**
* **Search:** Full-text search across all note titles and content
* **Project filter:** Scope the note list to a specific project context
* **New Folder:** Create a named folder to organize notes by project, target, or research theme
* **Publish:** Make a note or folder publicly accessible via a shareable link, enabling broader dissemination of research summaries, protocol documentation, or findings
## Workflow
Notes are organized into folders for structured knowledge management. A folder might correspond to a project (e.g., "EGFR Program"), a target (e.g., "CDK4/6"), or a campaign phase (e.g., "Lead Optimization Q2"). Within each folder, individual notes capture specific findings, decisions, literature summaries, or experimental observations.
The Publish function enables teams to convert internal notes into accessible documentation shared with external collaborators, project stakeholders, or the broader scientific community without requiring platform access.
## Running the Engine
### Inputs
| Action | Description |
| ---------- | ----------------------------------------------------- |
| New folder | Create a folder to organize notes by project or theme |
| New note | Create a structured note within a folder |
| Share | Share a note or folder with named team members |
| Publish | Generate a public link for external access |
### Outputs
* **Organized note library:** Searchable, folder-structured repository of all research notes
* **Shared notes:** Notes accessible to invited team members under their Shared Notes view
* **Published documents:** Public or link-accessible versions of selected notes for external sharing
# RevOptimization
Source: https://docs.revilico.bio/docs/revoptimization
Utilizing Revilico’s Molecular Optimization Engine to Transform Lead Compounds to Optimized Candidates
## Why Use this product?
Revilico’s Molecular Optimization engine aims to generate a compound library from a starting molecule that is optimized for the properties that the user prioritizes. This engine will transform lead compounds into optimized drug candidates through AI guided reinforcement learning that simultaneously balances multiple design objectives, potency selectivity, ADME properties, and synthesizability. For Multi-parameter optimization, reinforcement learning serves as a great algorithmic solution to optimizing leads for several different properties simultaneously.

## Background
Revilico’s Molecular Optimization Engine has relatively the same backbone as the De Novo Library Generation engine, where it is based on a chemical language model where SMILES strings are sentences, and each token is a part of a SMILES string and is a certain word. To recap how this model works is that we first have a start token, this is a small fragment of our SMILES string. Using this
$$
P(\text{next token} \mid \text{previous tokens})
$$
We will construct the probability distribution of the next token or next part of the SMILES string based on the current string we have and the objectives we are trying to solve for across each generative step. We will then sample from the distribution and append it to the previous tokens and repeat this process until we reach our end token which is optimized towards a specific objective, scored by several engines that assess physicochemical properties, activity, or other criterion.
Now what is the difference between De Novo Library Generation and Molecular Optimization? With Molecular Optimization we will need to define a scoring objective (i.e. what is one property we would like to optimize for), then we will need to reweight this objective in comparison to the other metrics in which the model is capturing. Essentially De Novo generation is just brainstorming the different molecules we can generate based off of a starting protein target, or just randomly generated molecules with certain optimized properties. With Optimization, we are sampling, scoring, then updating the generator so it will produce better molecules in line with our scoring objective, across each iteration of reinforcement learning.
Now let us dive into the full workflow. First our scoring function can be defined as:
$$
\text{Score}(\text{molecule}) \rightarrow x \in [0,1]
$$
here X is some sort of scaled value from 0 (poor) to 1 (excellent). Things that can go into a score can be physiochemical property targets (e.g. MW in a range, LogP in a range, TPSA), similarity constraints using fingerprint similarity, penalties and hard constraints (e.g. remove toxic substructures, reactive groups), activity calculations using Revilico’s FEP suite, or other binding affinity scoring functions on the platform, and external predictors (e.g. ADMET AI outputs or solubility predictors).
Before diving into the algorithm, we need to understand these two pieces of terminology: prior and policy.The prior is the baseline generative model trained to reproduce chemically valid molecules with an unbiased distribution while the policy is the optimized version of the prior, fine-tuned towards specific objectives (biased distributions). The prior is the assumptions of the model’s parameters before seeing any input data or scoring objections. The policy is a map of the actions we should take which will lead to maximizing long term rewards(in terms of more optimized molecules).
We first start with our first De Novo Library Generation step. Here we can say that our policy = prior (i.e. we do not know anything about the data and therefore there are no biases in the model). Here we will generate our first set of molecules then score them. When we score we will be able to update our policy, by updating rewards and penalties for the steps based on whether it will generate preferred or undesirable molecules. As we go through more iterations within the reinforcement learning process, the policy will tighten depending on the scoring function, helping the generator to shift the prior towards a more biased chemical space, optimized for your specific properties. By assigning rewards and penalties we will have a token probability bias. To ensure that we do not keep doing the same path over and over and keep ourselves in a local maxima (the max reward of what we have seen so far, but not the max reward of the entire search space), we introduce a probability of exploring a different space rather than maximizing the total reward at that step to ensure molecular diversity. This will allow us to see if there is another path of generating a molecule that has a higher maximum reward than the current maximum. Note that it is important that we only do incremental changes during each iteration, allowing the model to learn which patterns actually improve performance over time. Essentially what we are doing is trying to balance exploitation of the reward with exploration of the search space (optimization of the given property while allowing the algorithm to have enough diverse exploration for novel hypotheses. We can define what optimization is trying to maximize with the following equation:
$$
L(\theta) = \mathbb{E}_{x \sim \pi_\theta}[R(x) - \beta \mathrm{KL}(\pi_\theta(x) \| \pi_{\text{prior}}(x))]
$$
Where $\pi_\theta(x)$ is the current policy, $\beta$ controls how strongly the policy is anchored to the prior, with small $\beta$ being aggressive optimization and large $\beta$ being safe exploration, and KL being the regularization term that penalizes drifting too far from known chemistry.
Based on the score we generate we will be able to update our policy and repeat this process of optimization over and over again until we reach our panel of molecules.
# RevOrbitals
Source: https://docs.revilico.bio/docs/revorbitals
A Crash Course on Revilico’s Molecular Orbital Analysis Engine
**Why Use this Engine?**\
Molecular Orbital Analysis provides early insight into a compound’s electronic stability, reactivity, and binding potential before costly simulations or experiments are run. By identifying where a molecule donates and accepts electrons, this tool helps prioritize drug-like candidates, reduce chemical risk, and guide rational optimization decisions early in discovery.
**Background**\
Electron stability is a key principle that determines how easily a molecule participates in unwanted chemical rations in biological and chemical environments, which impacts safety, metabolism, and robustness. A drug candidate must survive different environments such as the bloodstream, metabolic enzymes, and off-target biomolecules, in which electronic stability will be greatly tested. Revilico’s Molecular Orbital Analysis Engine proves to be an effective tool where a user will be able to screen for electron stability and reactivity in a high throughput manner which will have a strong influence in assessing chemical stability, metabolic liability risk, redox / charge-transfer tendencies, and SAR interpretation across analogs.
Before diving into the workflow, we need to have a basic understanding of Density Functional Theory (DFT). A general explanation of what DFT is trying to solve is that we want to understand where electrons are in a molecule (probability distribution of the electron field) and how strongly they are held, partly because electrons determine bonding, reactivity, stability, and charge transfer. The main issue is that with the number of electron interactions we have to track, it becomes almost impossible even for moderately sized molecules to calculate these interactions 1 by 1 using Schrodinger’s equation. The solution is that with DFT, we will model the molecule using the overall electron cloud (electron density) instead of accounting for each electron as a point calculation. Diving deeper in the theory DFT is based on solving the many-electron Schrodinger equation denoted by the following equation:
Where the equation will scale catastrophically (O(n!) where n is the number of atoms) with the number of electrons and will become computationally intractable for real drug-like molecules. The DFT solution will replace the many-electron wavefunction (Psi) with the electron density p(r) which depends on three spatial coordinates, regardless of system size. The total energy is written as a functional of this density:
$$
E[p] = T[p] + V_{\text{ext}}[p] + J[p] + E_{xc}[p]
$$
Where T\[p] is the approximated kinetic energy, V\_ext\[p] is the electron-nuclear attraction, J\[p] is the classical electron-electron repulsion, and E\_xc\[p] is the exchange correlation energy. In practice DFT is solved using Kohn-Sham orbitals which satisfy:
$$
\left[ \frac{1}{2} \nabla^2 + V_{\text{eff}}(r) \right] \phi_i(r) = \epsilon_i \phi_i(r)
$$
Electron density is reconstructed as
$$
p(r) = \sum_i^{occ} |\phi_i(r)|^2
$$
When running the engine, we will convert our SMILES string and convert them into 3D molecular structures. We will then run geometry optimization where we will relax the molecule into a low energy structure so artificial strain does not distort the energy calculations. You can think of this as a calibration or preprocessing step. We will then run our DFT calculations which will determine the electron energy, a set of molecular orbitals, and the energy of each orbital. From these calculations we can identify our HOMO (highest-energy orbital that is occupied) and LUMO (lowest-energy orbital that is unoccupied). Calculating the difference between our HOMO and LUMO will measure the minimum energy cost to promote an electron from the most weakly held occupied state to the most empty state.
## Interactive HOMO-LUMO Viewer
Explore molecular orbital analysis results in an interactive 3D viewer. Toggle HOMO and LUMO orbital isosurfaces, adjust opacity and isosurface levels, and view orbital energies and HOMO-LUMO gap values.
* **Orbital Visualization** — View HOMO (red/pink) and LUMO (blue) isosurfaces overlaid on the molecular structure.
* **Interactive Controls** — Toggle orbitals on/off, adjust opacity and isosurface level for detailed exploration.
* **Energy Analysis** — Review HOMO energy, LUMO energy, and HOMO-LUMO gap for each molecule.
# Pathway Activation
Source: https://docs.revilico.bio/docs/revpathway-activation
Signaling Pathway Activity Scoring and Fold-Change Prediction from Drug-Induced Transcriptional Stress
## Why Use This Engine?
In the documentation below, we will use Revilico's Pathway Activation assay engine to predict which oncogenic and stress-response signaling pathways are activated or suppressed by a drug across a cancer cell line panel. This assay provides a transcriptional-level view of the drug's mechanism of action, identifying whether the compound engages DNA damage response, apoptotic signaling, survival pathways, or developmental programs, and at what concentrations these pathway changes occur.
## Background
Cancer cells are driven by dysregulated signaling pathways that control proliferation, survival, and differentiation. When a drug perturbs these pathways, it triggers compensatory responses: survival pathways may be upregulated, apoptotic programs activated, and developmental pathways suppressed. Understanding which pathways are affected and in which direction helps predict both the therapeutic mechanism and the potential for adaptive resistance.
RevAssay models eight core oncology signaling pathways drawn from KEGG and Reactome, covering the major kinase cascades relevant to solid tumor biology. The pathway database contains 47 pathways with curated KEGG IDs, Reactome IDs, gene targets, and confidence scores derived from literature concordance. Pathway fold-changes are derived from the cellular stress signal of the GNN+MLP viability model and calibrated to reflect the directional biology of stress-response signaling.
## Simulation Model
For each pathway $p$, a direction (activated or inhibited) and a maximum fold-change scale factor $\alpha_p$ are defined:
**Activated pathways** (fold-change above 1):
$$
\text{FC}_p = 1 + s \cdot (\alpha_p - 1) \cdot \eta_p, \quad \eta_p \sim \mathcal{LN}(0, 0.25^2)
$$
**Inhibited pathways** (fold-change below 1):
$$
\text{FC}_p = 1 - s \cdot (1 - \alpha_p) \cdot \eta_p, \quad \eta_p \sim \mathcal{LN}(0, 0.20^2)
$$
Where $s = 1 - v(c)$ is the cellular stress and $\alpha_p$ is the pathway-specific saturation fold-change. A dimensionless pathway activity score normalized to the range \[0, 1] is derived from the fold-change:
$$
\text{score}_p = \min\left(1, \frac{|\log_2(\text{FC}_p)|}{3}\right)
$$
A score of 0 indicates no change from baseline; a score of 1 corresponds to at least an 8-fold change in either direction.
**Modeled pathways and expected behavior under cytotoxic stress:**
| Pathway | Direction | Max FC | Biological rationale |
| ---------------- | --------- | ------ | --------------------------------------------------------- |
| RAS/MAPK/ERK | Activated | 2.6x | Paradoxical ERK reactivation under EGFR blockade |
| NF-kB | Activated | 3.2x | Master regulator of stress and survival signaling |
| p53 | Activated | 5.0x | DNA damage sensor; apoptosis trigger |
| JAK/STAT | Activated | 2.0x | Cytokine-driven survival program |
| PI3K/AKT/mTOR | Activated | 1.8x | Survival signaling (biphasic response) |
| Wnt/beta-catenin | Inhibited | 0.35x | EGFR-Wnt axis suppressed with EGFR blockade |
| Notch | Inhibited | 0.45x | Growth-promoting developmental pathway; stress-suppressed |
| Hedgehog | Inhibited | 0.50x | Developmental pathway; suppressed under cytotoxic stress |
## Parameters
| Parameter | Default | Description |
| ---------------------- | ------- | -------------------------- |
| Pinned pathways | All 8 | Which pathways to display |
| Exposure hours | 72 h | Duration of drug treatment |
| Hill coefficient | 1.2 | Dose-response steepness |
| Biological variability | None | Noise level for replicates |
## Outputs
* **Pathway fold-changes:** Per-pathway FC at each drug concentration for each cell line
* **Pathway activity scores:** Normalized \[0, 1] score per pathway
* **Direction classification:** Activated (red) or inhibited (blue) per pathway per concentration
* **Horizontal bar chart:** Pathways ranked by fold-change magnitude, colored by direction
* **Dose-response curves:** Per-pathway fold-change across the concentration range
* **96-well heatmap:** Plate-view colored by selected pathway activity score
* **Pathway metadata:** KEGG ID, Reactome ID, gene targets, and confidence score for each modeled pathway
# RevPhore
Source: https://docs.revilico.bio/docs/revphore
Understanding Interaction Dynamics using Revilico’s Pharmacophore Analysis Engine
## Why Use this product?
Revilico’s Pharmacophore Analysis engine is a tool that can automatically detect and map critical binding features, generate 2D interaction diagrams, comprehensive text reports, and interactive 3D visualizations that reveal the pharmacophoric requirements for target binding and guide lead optimization strategies. This engine is best used when you need to visualize and analyze protein-ligand binding interactions to understand molecular recognition patterns for structure-based drug design.Utilizing this engine paints a picture of binding mapping within the protein’s pharmacophore and allows for chemists to make more informed decisions on medicinal chemistry transformations they’d like to optimize for in subsequent lead optimization campaigns.
**Background**\
A pharmacophore defines binding interactions between each atom of a protein and ligand, and represents the minimum set of chemical features and binding modalities a molecule exhibits to bind well to a target. Pharmacophore nad binding analysis also can help to match the behavior of known active molecules with defined interaction sets. Along with focusing on exact atom-atom interactions, it focuses on defining roles across the engineered molecule like hydrogen-bond donor, hydrogen-bond acceptor, hydrophobic region, aromatic ring, and positive or negative charges. Essentially many different molecules can share similar pharmacophores if they present the same interaction features in roughly the same 3D arrangements. Pharmacophores are used to help chemists find new scaffolds that still fit the binding site, and to allow for chemists to use that information to make well defined alterations of their leads to optimize the interactions for the protein pocket and amino acids of interest.
Now how does it work? We take our protein-ligand complex defined and calculated by our virtual screening engine or co-folding algorithms and convert the molecules into decomposed 3D structures. Here the engine is able to identify interaction features on each structure (i.e. H-bond donors / acceptors, aromatic center, hydrophobic regions, and charged centers). It does this by parsing the inputted file for atoms and bonds, where it will apply chemical feature rules (i.e. N-H or O-H groups are considered H-bond donors). Then each feature is given a position in space (i.e. a single atom is given atom coordinates, a ring is given a ring centroid, and a functional group is given a geometric center). It will also check which ligand features are within interaction distance of protein residues to ensure physically rational assessments. This will confirm H-bond geometries, salt bridges, π - π stacking, and hydrophobic contacts. This will result in a set of labeled points in 3D space representing our ligand interaction map.
The next step is alignment and selecting key pharmacophore features. For this, the model will capture only the interactions that stabilize binding, not every possible feature the ligand has. We have our panel of features, and we can then filter down this panel of features to relevant interactions and binding patterns primarily based on distance (e.g. whether the donor or acceptor is within hydrogen-bonding range), orientation (e.g. whether the angle is reasonable for a H-bond), environment (e.g. whether the hydrophobic feature is near hydrophobic residues), and redundancy (e.g. whether multiple features are describing the same interaction). We then conduct feature merging, where we merge multiple atoms contributing to the same type of interaction into one interaction (e.g. a phenyl ring that has many carbons can be merged to form one aromatic interaction to reduce redundancies in our reporting).
In the end we will return a set of engagement patterns in 3D space where each interaction has a feature type (i.e. hydrogen-bond donor, hydrogen-bond acceptor, hydrophobic region, aromatic ring, positive or negative charge), 3D coordinates (where the interaction occurs in space), and tolerance radius (how much flexibility is allowed around that position). The Pharmacophore engine defines what the interaction is, where the interaction is taking place, and how much flexibility we have in that particular location.
**Credit to Revilico Advisor Pritam Kumar Panda:**
Pritam Kumar Panda. (2025). Protein AND ligAnd interaction MAPper: A Python package for visualizing protein-ligand interactions with 2D ligand structure representation. GitHub repository. [https://github.com/pritampanda15/PandaMap](https://github.com/pritampanda15/PandaMap)
## Interactive Pharmacophore Viewer
Explore pharmacophore analysis results in an interactive viewer. View 3D protein-ligand structures, 2D pharmacophore interaction maps, and detailed interaction breakdowns including hydrogen bonds, hydrophobic contacts, and aromatic interactions.
# RevPhospho
Source: https://docs.revilico.bio/docs/revphospho
Kinase Activity Inference and Drug Target Prioritization from Phosphoproteomics Data
## Why Use This Engine?
In the documentation below, we will use Revilico's RevPhospho engine to analyze phosphoproteomics data, infer kinase activity from phosphosite changes, map samples against cancer subtype signatures, and prioritize druggable kinase targets. This engine bridges the gap between raw mass spectrometry output and actionable drug discovery hypotheses by linking observed phosphorylation changes to the upstream kinases driving them.
## Background
Post-translational modifications (PTMs) are chemical modifications to proteins that occur after translation, dramatically expanding the functional diversity of the proteome beyond what is encoded in the genome. Phosphorylation, the addition of a phosphate group to a serine, threonine, or tyrosine residue by a kinase, is the most extensively studied PTM. It acts as a molecular switch that controls protein activity, localization, stability, and protein-protein interactions, making kinases among the most important and frequently targeted protein families in drug discovery.
Phosphoproteomics is the large-scale measurement of phosphorylation events across the proteome using mass spectrometry. It produces lists of quantified phosphosites with associated fold-change values between experimental conditions, providing a snapshot of signaling network activity. However, the direct interpretation of thousands of individual site-level changes is challenging. RevPhospho addresses this by using kinase-substrate databases to aggregate site-level signals into per-kinase activity scores, identifying which kinases are most active in the sample, comparing these activity profiles against cancer subtype reference signatures, and ranking kinases by their combined activity, druggability, and clinical evidence.
## Input Formats and Data Loading
RevPhospho accepts three input formats, auto-detected or specified by the user.
**MaxQuant format** is the output from the MaxQuant proteomics software suite. The engine reads columns for gene names, amino acid, position, localization probability, and normalized ratio values. Reverse database hits and common contaminants (marked with "+" in the corresponding columns) are filtered out before analysis. Phosphosites with a localization probability below 0.75 are excluded as their assignment to a specific residue is insufficiently confident.
**Generic CSV/TSV format** accepts any tabular file with columns for gene symbol, site (e.g. S473), log2 fold-change, and p-value. The separator (comma or tab) is auto-detected.
**PhosphoSitePlus format** accepts `.txt` exports from the PhosphoSitePlus database with gene and modification residue columns (e.g. S15-p) and header comment lines.
All three formats are normalized into a common internal representation with fields for gene symbol, site identifier, PTM type, PTM code, amino acid, log2 fold-change, p-value, and localization score. PTM types supported are phosphorylation (p), acetylation (ac), methylation (me), and ubiquitination (ub).
## Quality Control
Before any analysis, the engine computes the following QC metrics from the parsed input: total number of sites in the file, number of sites with a valid quantified log2 fold-change value, number of unique protein-coding genes represented, count of significantly upregulated sites (log2FC greater than 0.5), count of significantly downregulated sites (log2FC less than -0.5), and median absolute log2 fold-change across all quantified sites. These statistics are reported in the QC summary and flagged if the data appears sparse or has an unusual fold-change distribution.
## Kinase Substrate Enrichment Analysis
Kinase Substrate Enrichment Analysis (KSEA) is the core algorithm of RevPhospho. It translates site-level phosphorylation fold-changes into kinase-level activity scores by aggregating the fold-changes of all known substrates of each kinase and testing whether they are collectively shifted relative to the background distribution of all measured sites.
**KSEA Z-Score**
For each kinase $k$, the activity Z-score is computed as:
$$
Z_k = \frac{\overline{\text{FC}}_k - \overline{\text{FC}}_{\text{all}}}{\sigma_{\text{all}}} \cdot \sqrt{n_k}
$$
Where $\overline{\text{FC}}_k$ is the mean log2 fold-change of the quantified substrates of kinase $k$, $\overline{\text{FC}}_{\text{all}}$ is the mean log2 fold-change across all quantified phosphosites, $\sigma_{\text{all}}$ is the standard deviation of log2 fold-changes across all sites, and $n_k$ is the number of quantified substrates matched for kinase $k$. A positive Z-score indicates that the kinase's substrates are collectively upregulated relative to the background, implying kinase activation. A negative Z-score indicates that substrates are collectively downregulated, implying kinase inhibition or reduced activity.
**Statistical Testing**
In addition to the Z-score, a one-sample t-test is applied per kinase to assess whether the distribution of substrate fold-changes is significantly different from the global mean. All resulting p-values are corrected for multiple testing using the Benjamini-Hochberg false discovery rate procedure with a default significance threshold of FDR 0.05. Kinases with fewer than `min_substrates` matched quantified substrates (default 3) are excluded from reporting to prevent low-confidence estimates from sparse substrate coverage.
The kinase-substrate reference database integrates annotations from PhosphoSitePlus, NetworKIN, and SIGNOR, covering over 150 human kinases. Each kinase entry includes its gene symbol, full name, kinase family (RTK, SFK, AGC, PIKK, CMGC, and others), and its curated list of known substrates.
## Cancer Signature Scoring
Beyond identifying which kinases are active, RevPhospho maps the sample's kinase activity profile against 12 reference cancer subtype signatures derived from CPTAC (Clinical Proteomic Tumor Analysis Consortium) phosphoproteomic datasets. This enables researchers to determine which cancer subtype the sample's signaling landscape most closely resembles.
**Cosine Similarity**
For each cancer signature $s$, the similarity to the sample's kinase Z-score vector is computed using cosine similarity:
$$
\text{similarity}_s = \frac{\mathbf{z} \cdot \mathbf{v}_s}{\|\mathbf{z}\| \cdot \|\mathbf{v}_s\|}
$$
Where $\mathbf{z}$ is the vector of kinase Z-scores from the sample and $\mathbf{v}_s$ is the reference activity direction vector for signature $s$. The similarity is normalized to a \[0, 1] score:
$$
\text{score}_s = \frac{\text{similarity}_s + 1}{2}
$$
Scores are classified into three confidence levels: High (score above 0.65), Medium (0.45 to 0.65), and Low (below 0.45). The 12 signatures span five cancer types across their major molecular subtypes:
| Cancer Type | Subtypes |
| -------------------------- | ------------------------------------------------- |
| Breast (BRCA) | HER2-enriched, Luminal A, Triple-Negative |
| Lung Adenocarcinoma (LUAD) | EGFR-mutant, KRAS-mutant, ALK/ROS1-fusion |
| Colorectal (COAD) | Chromosomal Instability, MSI-Hypermutated |
| Glioblastoma (GBM) | Classical (EGFR-amplified), Mesenchymal |
| Uterine (UCEC) | Copy-Number High (Serous-like), POLE-ultramutated |
## Drug Target Prioritization
The drug target prioritization module combines the kinase activity Z-score, curated druggability, and cancer relevance into a single ranked score to identify the most actionable therapeutic targets in the sample.
**Combined Score**
$$
\text{score}_k = \frac{|Z_k|}{Z_{\max}} \cdot 4.0 + D_k \cdot 4.0 + R_k \cdot 2.0
$$
Where $|Z_k| / Z_{\max}$ is the normalized absolute activity score (scaled to 0-4), $D_k$ is the curated druggability score for kinase $k$ on a scale of 0 to 1 (reflecting the existence of approved drugs, clinical-stage compounds, or published tool compounds), and $R_k$ is the cancer relevance weight derived from the cancer signature scores (1.0 for high confidence match, 0.5 for medium, 0.1 for low). Kinases with negative Z-scores (inhibited kinases) receive a 15% score penalty, as activated kinases are generally higher-priority therapeutic targets than inhibited ones. The top 25 kinases by combined score are returned as the prioritized target list.
## PTM Site Annotation
Each phosphosite quantified in the input is cross-referenced against a curated annotation database covering over 50 high-value cancer-relevant phosphosites across 10 key oncoproteins including TP53, AKT1, EGFR, SRC, BRCA1, CTNNB1, MYC, STAT3, CDK2, and MDM2. For each annotated site, the engine reports the known biological function (e.g. activation loop phosphorylation, degron phosphorylation), the known effect (activating, inhibitory, modulatory, or DNA damage response), the kinases known to phosphorylate that site, and associated disease contexts. Sites not present in the curated database are passed through without annotation.
## PTM Crosstalk Detection
The crosstalk detection module identifies regulatory relationships between pairs of phosphosites on the same protein that are both measured in the input dataset. Crosstalk events are classified into four mechanistic types.
**Priming** occurs when phosphorylation of site A creates a recognition motif that enables a second kinase to phosphorylate site B. For example, CK1-mediated phosphorylation of CTNNB1 S45 creates the priming event that allows GSK3 to phosphorylate S33, S37, and T41 in the beta-catenin destruction complex.
**Cooperative** events occur when co-phosphorylation of both sites A and B is required for full protein activation. For example, AKT1 requires phosphorylation at both T308 (by PDK1) and S473 (by mTORC2) for maximal kinase activity.
**Competitive** events occur when sites A and B compete for the same binding domain or reader protein, such that phosphorylation of one site can displace binding at the other.
**Reader** events occur when phosphorylation of site A recruits a reader protein that then modifies or regulates the phosphorylation state of site B.
The crosstalk database covers over 20 curated events. For each event where both sites are present in the input data, the engine reports the protein, both site identifiers, the crosstalk type, a mechanistic description, supporting literature, and the observed log2 fold-change at each site.
## Protein Search and Structure Visualization
In addition to the phosphoproteomics analysis pipeline, RevPhospho provides a protein-centric search view. Users can search for any human protein by gene symbol, protein name, or synonym. For proteins in the curated database, the engine returns all annotated PTM sites with residue-level detail including known kinases, functional descriptions, literature support counts from low-throughput and high-throughput studies, and disease associations. For proteins outside the curated set, annotations are retrieved from UniProt via the public API. Protein 3D structures are rendered interactively using AlphaFold v4 structures retrieved from the EBI API via the NGL Viewer.
## Running the Engine
### Inputs
| Parameter | Default | Description |
| ---------------- | -------- | ---------------------------------------------------------------------- |
| `ptm_file` | Optional | Phosphoproteomics file in MaxQuant, Generic, or PhosphoSitePlus format |
| `input_format` | maxquant | File format: `maxquant`, `generic`, or `phosphositePlus` |
| `organism` | human | `human` or `mouse` |
| `min_substrates` | 3 | Minimum matched substrates per kinase for KSEA inclusion |
| `ptm_types` | p | Comma-separated PTM codes to analyze: `p`, `ac`, `me`, `ub` |
A built-in demo dataset simulating a lung adenocarcinoma phosphoproteome is available for immediate exploration without uploading a file.
### Outputs
Upon completion, the engine returns results across six modules:
* **QC summary:** Site counts, protein counts, upregulated and downregulated site counts, median absolute fold-change
* **KSEA results:** Per-kinase Z-score, substrate fold-change mean, matched substrate count, p-value and FDR, kinase family, and list of matched substrate site IDs
* **Cancer signatures:** Per-signature similarity score and confidence level across all 12 CPTAC-derived subtypes
* **Drug targets:** Top 25 ranked kinases with combined priority score, druggability, cancer relevance weight, and Z-score component
* **PTM site annotations:** Functional annotations, known effect, kinase assignments, and disease associations for each recognized site
* **Crosstalk events:** Detected inter-site regulatory relationships with mechanistic classification and both site fold-changes
# RevPocket
Source: https://docs.revilico.bio/docs/revpocket
## Why use this engine?
In the documentation below, we will use Revilico’s Pocket Search Engine to identify potential pockets to target with drug binding. Using Pocket search we will identify pockets on the static version of the protein based simply on the geometry of the protein and use MDPocket to simulate which pockets remain stable while the protein is in motion. This portion of computational chemistry is essential to understanding how key biological functions can be enhanced or blocked using different therapeutic strategies. This algorithm mainly searches for druggability and chemistry, and when supplemented with prior biological knowledge and literature review, it can be a powerful tool to regulate biologies.

## Background
Pocket identification is a critical aspect of drug discovery. It allows researchers to understand protein function and enable rational drug design by pinpointing specific areas along the protein where small molecules, ligands, or other proteins can bind. In Revilico’s Pocket Search Engine, we introduce two tools, one for understanding the protein from a static perspective and one for understanding the protein from a dynamic perspective. Our pocket search engine is a purely geometry based algorithm that detects binding pockets by analyzing protein shape and cavity architecture.
We will now dive into the theory behind the algorithm on which the system was built on. We first start with our protein. On each atom of this protein, we assign an atomic exclusion zone defined by its van der Waals (dW) radius. The circumference of this circle defines the boundaries of this atom, and for each element there will be an assigned radius in Angstroms where no ligand atom can exist. We then define all other space not in this atomic exclusion zone as potential void space in which a ligand atom can exist. The next step is understanding within this void where we can identify these void centers. This void center identification is based on Voronoi tessellation which is a mathematical framework defined by the following equation:
For a set of protein atoms $\{A_1, A_2, \ldots, A_n\}$, the Voronoi cell of atom $A_i$ is
$$
V(A_i) = p \in \mathbb{R}^3 : \|p - A_i\| \le \|p - A_j\| \text{ for all } j \ne i
$$
Where p is any point in 3D space. Essentially what this means is that a Voronoi cell contains all points closer to Ai than to any other atom. Where we identify a void center is that this is a point that is equidistant from typically 4 or sometimes more voronoi centers. We then generate a sphere with the center located at this void center, and expand until it touches the surface of the nearest atom. If the sphere generated is greater than our min alpha sphere radius and less than our max alpha sphere radius. It remains a valid alpha sphere. The min and max of the sphere sizes are pre-determined. Once multiple valid spheres are generated, if we meet our min spheres per pocket parameter within our clustering distance, it is then classified as a pocket. Another way to look at this equation and system is that we are able to ‘roll a sphere’ across the entire protein and detect the chemistries along that sphere’s route to be able to extract critical parameters that we can use in downstream assessments.
**Identifying Pockets with Revilico’s Pocket Search Engine**\
Pocket search is an engine where the primary goal is to answer the question: Where are the physical cavities large enough to accommodate drug-like molecules. This engine is essential for identifying potential binding sites for drug design, rank pockets by druggability before expensive docking calculations, mapping surface topology for understanding protein function, selecting docking regions based on geometric feasibility, and generating hypotheses for mutagenesis or fragment screening. Usually, you can use literature based on experimental methods to determine regions on the amino acid chains that contribute to biological activity to overlay with these pocket search algorithm regions to confirm or identify pockets.
**Identifying Pockets with Revilico’s MDPocket Engine**\
MD Pocket operates the same way as Pocket search, the difference being that it calculates the Pocket search at every frame of a MD simulation, capturing the pockets while the protein dynamically changes, compared to Pocket search which just captures the pockets at one screenshot of the protein. MDPocket reveals which pockets persist, disappear, or transiently form over time.
## Interactive Pocket Viewer
Explore FPocket pocket detection results in an interactive 3D viewer. The protein structure is shown at 50% transparency with detected pockets highlighted as spacefill overlays.
* **All Pockets View** — See all detected pockets overlaid in red on the protein structure.
* **Individual Pocket** — Select a specific pocket to highlight it in orange with detailed statistics.
* **Pocket Statistics** — View druggability score, volume, SASA, hydrophobicity, and more for each pocket.
# Proliferation (GR)
Source: https://docs.revilico.bio/docs/revproliferation-gr
Growth Rate-Corrected Proliferation Analysis Separating Cytostasis from Cytotoxicity
## Why Use This Engine?
In the documentation below, we will use Revilico's Proliferation (GR) assay engine to compute growth rate-corrected drug sensitivity metrics that distinguish whether a compound is slowing cell division (cytostatic) or actively killing cells (cytotoxic). Standard IC50 measurements confound these two fundamentally different drug effects. The GR correction method (Hafner et al., 2016) removes the dependence of apparent drug potency on the cell line's intrinsic proliferation rate, providing more consistent and mechanistically interpretable metrics across cell lines with different doubling times.
## Background
A drug that arrests cells in G1 without killing them will appear potent in a standard endpoint viability assay simply because cells in a rapidly dividing control well have doubled or tripled in number while treated cells have not. The measured IC50 is therefore a function of both the drug's mechanism and the cell line's proliferation rate, making cross-cell-line comparisons unreliable. The GR correction method accounts for this by normalizing the measured cell count response to the expected cell count in the absence of drug, based on the known cell doubling time.
The GR value maps the response to a scale where +1 indicates no drug effect (full growth), 0 indicates cytostasis (alive but not dividing), and -1 indicates a net cytotoxic effect (more cells dying than dividing). This scale is mechanistically interpretable: drugs with GR\_max near 0 are pure cytostatic agents (kinase inhibitors at therapeutic doses), while drugs with GR\_max near -1 are cytotoxic agents (chemotherapies).
## Simulation Model
The GR value at concentration $c$ is computed directly from the GNN+MLP viability prediction $v(c)$:
$$
\text{GR}(c) = 2 \cdot v(c) - 1 + \varepsilon, \quad \varepsilon \sim \mathcal{N}(0, 0.08^2), \quad \text{clamped to } [-1, +1]
$$
This linear mapping reflects the mathematical relationship between normalized viability and GR: a viability of 1.0 (no effect) maps to GR = +1, viability of 0.5 maps to GR = 0 (cytostasis), and viability of 0 maps to GR = -1 (complete cytotoxicity).
**GR50** is the concentration at which GR = 0.5 (half-maximal growth inhibition), extracted from the GR dose-response curve and analogous to IC50 but corrected for proliferation rate.
**Doubling time** is modeled as increasing with drug stress:
$$
\tau(c) = \frac{\tau_0}{\max(v(c), 0.05)} \cdot \eta_\tau, \quad \eta_\tau \sim \mathcal{LN}(0, 0.10^2)
$$
**Basal doubling times** by cell line:
| Cell Line | Doubling time | Proliferation rate |
| ---------- | ------------- | ------------------------- |
| MDA-MB-231 | 20 h | Fast (aggressive TNBC) |
| A549 | 22 h | Fast (NSCLC) |
| MCF7 | 28 h | Moderate (luminal breast) |
| HEPG2 | 32 h | Slow (hepatocellular) |
**EdU incorporation** (fraction of cells in S-phase, measured by 5-ethynyl-2'-deoxyuridine incorporation as a proxy for active DNA synthesis):
$$
\text{EdU\%}(c) = v(c) \times 55 + \varepsilon_{\text{EdU}}, \quad \varepsilon_{\text{EdU}} \sim \mathcal{N}(0, 6^2)
$$
The scaling factor of 55% reflects the average S-phase fraction of cycling cancer cells in an asynchronous population.
## GR Metric Interpretation
| GR value | Interpretation | Drug behavior |
| -------- | ------------------------------------- | ------------------------------ |
| +1 | No effect | Full growth, drug inactive |
| +0.5 | Half-maximal growth inhibition (GR50) | Partial cytostasis |
| 0 | Complete cytostasis | Cells alive but not dividing |
| -0.5 | Partial cytotoxicity | Net cell loss at moderate rate |
| -1 | Complete cytotoxicity | All cells dying |
## Parameters
| Parameter | Default | Description |
| ---------------------- | ------- | ------------------------------------- |
| EdU detection | On | Enable S-phase fraction measurement |
| GR metrics | On | Enable GR value and GR50 computation |
| Assay days | 3 | Duration context for GR normalization |
| Exposure hours | 72 h | Drug treatment duration |
| Hill coefficient | 1.2 | Dose-response steepness |
| Biological variability | None | Noise level for replicates |
## Outputs
* **GR value:** Growth rate-corrected drug effect at each concentration per cell line (range -1 to +1)
* **GR50:** Concentration at half-maximal growth inhibition per cell line
* **GR\_max:** GR value at saturating drug concentration (cytotoxicity ceiling)
* **Doubling time:** Drug-extended cell cycle duration per cell line per concentration
* **EdU %:** Fraction of cells in S-phase as a DNA synthesis proxy
* **GR dose-response curves:** Per-cell-line GR vs. log concentration curves
* **Comparative GR plot:** All cell lines overlaid with GR = 0 cytostasis reference line
* **96-well heatmap:** Plate-view colored by GR value
* **GR50 waterfall:** Cell lines ranked by GR50, analogous to the IC50 waterfall in RevViability
# RevQS
Source: https://docs.revilico.bio/docs/revqs
DFT-Based Quantum Mechanical Scoring of Protein-Ligand Binding Interactions
## Why Use This Engine?
In the documentation below, we will use Revilico's RevQS engine to compute the quantum mechanical binding interaction energy between a ligand and its protein binding pocket using density functional theory (DFT). Unlike empirical scoring functions or machine learning models trained on binding data, RevQS derives the interaction energy directly from first principles, capturing electrostatics, exchange-repulsion, polarization, charge-transfer, and London dispersion from the electronic wavefunction. This provides a physics-grounded binding score that is particularly valuable for lead optimization, distinguishing close structural analogs, and validating binding hypotheses in cases where force-field or ML approaches give inconsistent results.
## Background
Classical molecular docking scoring functions approximate binding affinity using parametrized empirical terms. Machine learning scoring functions learn statistical patterns from experimental data. Both approaches are fast but rely on approximations that can fail for unusual chemistries, charged systems, or novel scaffolds outside the training distribution. DFT scoring computes the quantum mechanical wavefunction of the entire protein pocket and ligand system, capturing all physical interaction terms from first principles without empirical parametrization.
The central challenge in applying DFT to protein-ligand systems is the basis-set superposition error (BSSE), an artifactual stabilization that arises when the basis functions of one fragment assist the description of the other. RevQS corrects for this using the Boys-Bernardi counterpoise (CP) method, which computes the interaction energy using a consistent extended basis set that spans both fragments. The result is an interaction energy that accurately reflects the true quantum mechanical binding contribution.
## Pocket Preparation
Given a protein PDB file and a ligand SDF file containing a docked 3D pose, RevQS first prepares the binding pocket. Protonation states are assigned at the specified pH (default 7.4) using PDBFixer. All protein residues within a user-defined cutoff radius (default 10 Angstroms) of any ligand atom are included in the pocket fragment. If the number of atoms exceeds the maximum (default 120 for def2-SVP), the farthest residues are trimmed until the limit is met. Backbone bonds cut at the residue boundaries are capped with hydrogen atoms to maintain chemical valence. The total pocket charge is estimated from standard amino acid formal charges (ARG and LYS contribute +1, ASP and GLU contribute -1).
## DFT Calculation and CP Correction
Three separate single-point DFT calculations are performed to implement the Boys-Bernardi counterpoise correction:
**E(complex):** The full quantum mechanical calculation of the protein pocket and ligand together, using the combined basis set of all atoms.
**E(protein, ghost ligand):** The protein pocket calculation with the ligand atoms represented as ghost atoms, providing basis functions but no electron density or nuclear charges.
**E(ligand, ghost protein):** The ligand calculation with the protein pocket atoms as ghost atoms.
The counterpoise-corrected interaction energy is:
$$
\Delta E_{\text{CP}} = E(\text{complex}) - E(\text{protein}_{\text{ghost}}) - E(\text{ligand}_{\text{ghost}})
$$
This correction removes the artificial stabilization caused by basis set incompleteness. The RevQS score is reported as the negative of the interaction energy so that more favorable binding gives a larger positive score:
$$
\text{RevQS score} = -\Delta E_{\text{CP}} \quad (\text{kcal/mol})
$$
Ligand efficiency is computed as the RevQS score normalized by the number of heavy atoms:
$$
\text{LE} = \frac{\text{RevQS score}}{N_{\text{heavy atoms}}} \quad (\text{kcal/mol per heavy atom})
$$
## DFT Method
**Default functional: wB97M-V**
RevQS uses the range-separated hybrid meta-GGA functional wB97M-V with built-in VV10 non-local correlation. This functional was specifically parametrized on the MGCDB84 database to describe non-covalent interactions accurately, including London dispersion, polarization, and charge-transfer, without any additional empirical dispersion correction. It consistently outperforms B3LYP-D3 and other popular functionals on benchmark sets for protein-ligand non-covalent interactions.
**Default basis set: def2-SVP**
The def2-SVP double-zeta polarization basis set provides a good balance between accuracy and computational cost for pocket sizes of up to 250 atoms on a single A10G GPU (24 GB VRAM). For higher accuracy, def2-TZVP (triple-zeta) reduces basis-set incompleteness significantly but limits the system size to approximately 120 atoms on the same hardware.
**Alternative functionals** for faster calculations include B3LYP-D3(BJ), PBE0-D3(BJ), and M06-2X, which are 3 to 4 times faster than wB97M-V at some cost to accuracy for dispersion-dominated interactions.
GPU acceleration is provided via GPU4PySCF, which implements the DFT calculation on NVIDIA GPUs using a batched Coulomb and exchange matrix build algorithm. Calculations fall back to CPU automatically if no CUDA hardware is available.
## F-SAPT Decomposition
When the F-SAPT option is enabled, the interaction energy is further decomposed into four physically interpretable components using functional-group symmetry-adapted perturbation theory (F-SAPT) via Psi4:
$$
\Delta E_{\text{int}} = E_{\text{elec}} + E_{\text{exch}} + E_{\text{ind}} + E_{\text{disp}}
$$
Where $E_{\text{elec}}$ is the classical electrostatic interaction between unperturbed charge densities, $E_{\text{exch}}$ is the Pauli exchange-repulsion arising from wavefunction antisymmetry, $E_{\text{ind}}$ captures polarization and charge-transfer response, and $E_{\text{disp}}$ captures London dispersion forces from correlated electron fluctuations. This decomposition reveals whether binding is primarily electrostatically or dispersion driven, which directly informs medicinal chemistry decisions about which substituents to modify.
## Interaction Map
The interaction map decomposes the total interaction energy into per-residue contributions using a frozen-background DFT approximation, providing an atomic-resolution picture of which binding site residues contribute most to the overall RevQS score. Each residue's contribution is classified by dominant interaction type (electrostatic or dispersion), enabling structure-guided optimization decisions such as which residues to target for improved van der Waals contacts or hydrogen bond geometry.
## Running the Engine
### Inputs
| Parameter | Default | Description |
| ------------------- | -------- | ------------------------------------------ |
| Protein PDB | Required | Protein structure (ligand removed) |
| Ligand SDF | Required | Pre-docked 3D pose (protonated) |
| Functional | wb97m-v | DFT exchange-correlation functional |
| Basis set | def2-svp | Gaussian basis set |
| Pocket cutoff | 10.0 A | Residue inclusion radius around ligand |
| Max pocket atoms | 120 | Atom count cap for GPU memory management |
| pH | 7.4 | Target pH for protonation state assignment |
| Run F-SAPT | False | Enable interaction energy decomposition |
| Run interaction map | True | Enable per-residue contribution analysis |
### Outputs
* **RevQS score:** CP-corrected DFT interaction energy in kcal/mol (positive = favorable)
* **Ligand efficiency:** RevQS score per heavy atom in kcal/mol
* **BSSE correction:** Magnitude of the counterpoise correction in kcal/mol
* **F-SAPT decomposition:** Electrostatic, exchange, induction, and dispersion components (if enabled)
* **Interaction map:** Per-residue interaction energies with dominant interaction type classification
* **Pocket composition:** Number of residues, atoms, and charges in the pocket fragment
* **Computation metadata:** Functional, basis set, wall time, GPU used
# RevQSAR
Source: https://docs.revilico.bio/docs/revqsar
Utilizing Revilico’s Quantitative Structure Activity Relationship (QSAR) Model for Understanding Chemical Spaces
## Why Use this product?
After retrieving a variety of different data types from your engines across the entire platform, you will likely end up with a matrix of SMILES strings and data points which need downstream processing and analysis. One method of understanding chemical similarities within clusters of high performing compounds is through the development of QSAR models that help you to analyze your chemical space, extract substructures of interest, and determine what your ideal scaffolds to preserve should be before moving into lead expansions. The way this model works is by taking and ingesting data collected from across the platform, extracting chemical feature representations of your molecules, and running them through different clustering modalities to explore chemical space in a data driven, intuitive way before moving into generative chemistry or lead expansion.
## Background
Usually, when running something like a high throughput screen for collecting hits for your therapeutic campaign, you collect a lot of data that looks like a compound and its associated property values (for activity, solubility, etc.). Generally, what we can begin doing in later stage development flows is to take in this data for feature extraction and for analysis through clustering (which is a lower dimensional representation of chemical space). In other words, we can take a set of compounds and their determined properties and create a map of compounds that are grouped together based on how similar their structures are, and overlay the property of interest that you’d like to analyze to get a better understanding of what regions on the molecule, overlays, or substructures drive the certain properties that you are evaluating. In the case of activity, you usually have a portion of the molecule that is a critical driver of binding, and therefore needs to be preserved. With this technique, we can extract representative features from the SMILES structures themselves to understand what drives these properties.
When utilizing the algorithm, what is happening on the backend, is that the algorithm is ingesting your SMILES Strings along with numeric columns that represent your properties that you are interested in. Eventually, you will select your ‘featurization method’ which includes Morgan Fingerprints, MACCS Keys, GraphConv embeddings, or chemBERTA embeddings. Each embedding type provides specific benefits depending on the ‘task’ you are trying to predict and create representations about. For the above featurizers, Morgan Fingerprints (MFP) create extended connectivity fingerprints (ECFPs) that capture local topology and substructure patterns usually used for similarity searches, or QSAR. MACCS Keys are of fixed length and it decodes the structure of your molecule into a predefined SMARTS pattern (substructure like aromatic ring, carboxyl, etc) and is simple, but less expressive than MFP. GraphConv embeddings or Graph Convolutional Network embeddings (GCNs) are a deep learning representation where atoms are nodes and bonds are edges, helping to capture both local and global structural context where nonlinear relationships matter more. ChemBERTA embeddings are transformer based language models trained on SMILES Strings to learn contextual chemical representations, encoding semantic chemical language.
When selecting the clustering algorithm, it is important to note what your primary task is, what the nature of the chemical space is, and how you’d like to create lower dimensional representations for your chemical series. To begin, K-means clustering partitions molecules into k clusters by minimizing inter cluster variances and works best for continuous embeddings like GraphConv and chemBERTa rather than binary fingerprints like ECFP/MFP. Hierarchical clustering builds dendrograms by iteratively merging or splitting clusters based on distance metrics like tanimoto similarity or cosine similarities, without a pre-defined number of clusters(k), which is more common for fingerprint based similarity analysis. Spectral Clustering uses eigenvectors of the similarity matrix to perform clustering in a reduced ‘spectral’ space. This works best for non-spherical or complex cluster boundaries or anticipated continuous chemical spaces. Tanimoto based clustering is specialized for binary fingerprints like MFP/MACCS and uses the tanimoto similarity score to group molecules with shared substructure patterns and scaffolds.
Uniform Manifold Approximations and Projections (UMAP) and K-means is non-linear and preserves global and local structures and is best for larger chemical spaces of interest. T-distributed Stochastic Neighbor Embeddings (tSNE) and k mean emphasize local similarity and points close in high dimensional space remain close in the 2D and 3D projections, they are also non linear and emphasize local clusters and are best for visualization of the chemical space into distinct units. Principal Component Analysis (PCA) and k-means is a dimension reduction technique projecting data onto components that capture maximum variances, and is linear, fast, and an interpretable baseline for descriptor based clustering methods.
# RevRetro
Source: https://docs.revilico.bio/docs/revretro
Analyzing Synthesizability of Target Molecule Using Revilico’s Retrosynthesis Engine
## Why Use this product?
The Retrosynthesis engine is an AI powered tool that generates complete multi-step synthesis pathways, identifies required building blocks, and provides alternative routes ranked by feasibility enabling rapid route planning, synthesizability assessment, and identification of chemical analogs utilizing retrosynthetic analysis. This engine is best used when you have a target molecule and need to determine how to synthesize it from commercially available starting materials.

## Background
Synthesizability is a key metric in evaluating whether a molecule can become a drug in practice. Without synthesizability, even the most potent, selective, or computationally promising molecule cannot become practically made in the laboratory. In a laboratory, molecules are synthesized through a stepwise chemical synthesis starting from existing well characterized compounds that are drawn from commercial reagent catalogs, fragment libraries, and previously synthesized intermediates. Essentially, this compound library will be our building blocks for each of our molecules, and will be utilized in several reaction steps. When understanding synthesisability, we are asking the question of whether our molecule be designed using the current compound library that we have on hand? This is where Revilico’s Retrosynthesis Engine comes into place to help plan, organize, and arrange logistics for challenging synthesis steps of downstream lead series acquisition and testing. This engine has the ability to test this parameter and generate hypotheses for chemists from a computational standpoint, and in a high throughput manner, to enable chemists to thoroughly screen their library of lead molecules for synthesizability, saving both time and money.
Diving into the workflow of how this engine works, we realize that this is an advanced tree search problem, or in other words we are trying to find all possible combinations of building blocks needed to synthesize this molecule. We will first start with our molecule of interest and try to decompose it into the different building blocks (i.e. fragments, intermediates, and reagents) that can be curated into our target molecule of interest that needs to eventually be tested in the lab.
After designing a lead molecule computationally, the largest question will be how this compound is segmented into smaller substituents that can be readily accessible to chemists in the lab.Traditionally, this process would be handled primarily through the chemists experiences and theoretical recollection, but now we can begin to determine the building block combinations required to build this molecule. This model has been trained on experimental reaction data from a large library of known compounds including, commercially available reagents, common intermediates, and frequently used medicinal chemistry fragments where the input is reactant molecules and the output is the observed product. This relationship can be encoded by the following equation
$$
(\text{Reactant Set}) + (\text{Reaction}) \rightarrow (\text{Product})
$$
From the data the model will learn which reactions are chemically valid, which functional groups participate in which transformation, which precursor-reaction combinations are plausible,what the reaction conditions are, and how likely a given transformation is based on the precedent reaction sets encoded within the model.
Based on this foundation, we can dive deeper into Chain-of-Reaction (CoR) reasoning. Rather than predicting a single step, the model will generate CoR sequences where a CoR is an ordered sequence of building blocks and reaction steps. When executed in the forward direction, the chain will reconstruct the target molecule using prior knowledge of reactants and reaction sets within its training set. We can think of the entire Chain of Reaction sequence as a recipe, where each node will represent the progress in the recipe. We can think of each node as the step needed to create the final product, and with this model architecture, we are able to fragment the features that are fed in to represent a cohesive retrosynthesis problem. For example, one step is how we make the dough. Water, yeast, and flour are the reactants, mixing is the reaction and dough is the product. The CoR architectures are able to fully resolve these steps with different confidences to paint a picture of experimental protocols that should be taken to get to target products with on hand reactants.
Now we introduce the core algorithm. Our algorithm is based on a tree search, specifically tailored towards defining retrosynthesis as a decomposition problem. If we think about a retrosynthetic pathway as a tree/network with nodes and edges, we have defined the maximum number of child nodes that can be allocated per parent node, where a set of child nodes are the decomposition steps of the parent node. Nodes at the same hierarchy level represent different possible configurations to create the same parent node for a given task. We have also determined the depth of the ‘tree’ or network, which signifies how many times an ingredient can be decomposed. For example, if we are looking at baking bread, a depth of 3 could start with bread, decompose to dough, then decompose the dough into water, yeast, and flour, and finally decompose the flour into wheat.
We can now introduce the Beam Search algorithm. The Beam Search algorithm is a heuristic tree-search algorithm that explores many candidate solutions in parallel while strictly limiting how many are actively pursued at each step. In other words at each step we have a set of nodes and we will only take a few of the top performing ones and expand those. For example, say we have a search width of 3. First we start with our parent node and generate the top 3 children nodes. The next step will be expanding out each of these 3 nodes, so at the next level we will have 9 total nodes. At this step we will rescore and only take the top 3 at the same level again, then expand those, repeating the same steps until we hit our max depth. So essentially we start at 1 node, expand to 3, then expand to 9 only keeping the top 3 performing ones, then expanding to 9 again keeping the top 3, repeating over and over until we reach our depth length. We are searching for the best possible configurations or components that contribute to the decompositions.
Now how do we decompose each node? Each node is a set of different reactants and reactions. For now we will look at the reactants. In our bread analogy, say our dough is the current parent node. Then the children nodes or reactants will be our water, yeast and flour. The algorithm will look at each of these molecules or reactants in our set (water, yeast, and flour), then evaluate for the most difficult or unresolved molecule set for expansion. Molecules already in the building-block are deprioritized. So in our case the next level will be expanding flour and keeping water and yeast the same since those are building blocks, things that are elementary enough and do not need to be decomposed further. We can then decompose flour into grains. Now at this new node our reactants will be water, yeast, and grains.
In the end we should have a set or sets of reactants from our building-block library that can be used to synthesize our molecule. For the purpose of this algorithm, we utilize commercially available reactant sets primarily to ensure logistically smooth syntheses.
## Interactive Results Viewer
Explore retrosynthesis results interactively. View synthesis pathways ranked by feasibility score, reaction steps, similarity metrics, and 3D molecular structures.
# RevRisk
Source: https://docs.revilico.bio/docs/revrisk
Polygenic Risk Scoring Across 36 Disease Panels and the Full PGS Catalog from Whole-Genome VCF
## Why Use This Engine?
In the documentation below, we will use Revilico's RevRisk engine to compute polygenic risk scores (PRS) from a whole-genome VCF file across 36 built-in disease panels and any score from the PGS Catalog. This engine enables researchers and clinicians to stratify individual genetic risk across a broad range of complex diseases, compare scores against population-level reference distributions, and report standardized percentile rankings and relative risk estimates.
## Background
Polygenic risk scores aggregate the small effects of many common genetic variants identified through genome-wide association studies (GWAS) into a single summary statistic that estimates an individual's inherited predisposition to a given trait or disease. Each variant in a GWAS weight set contributes a dosage-weighted effect size to the total score. Because complex diseases are influenced by thousands of loci each with small individual effects, PRS captures the cumulative genetic burden that no single variant analysis can reveal.
PRS has emerged as a clinically relevant tool for identifying individuals at elevated lifetime risk of conditions such as coronary artery disease, type 2 diabetes, and several cancers, often providing predictive value that is independent of and complementary to conventional clinical risk factors. RevRisk implements a complete PRS pipeline: VCF parsing and quality control, dosage extraction, raw score computation, population-level standardization using Hardy-Weinberg equilibrium-derived statistics, conversion to percentile rank against a UK Biobank reference population, and risk category classification. The engine covers 36 curated built-in disease panels and integrates with the PGS Catalog to support scoring against any of over 4,000 published scores.
## VCF Parsing and Quality Control
The engine accepts per-sample genotype VCF files in `.vcf`, `.vcf.gz`, or `.bcf` format. After loading, each variant record passes through the following sequential quality control filters:
**PASS filter:** Variants with a FILTER field value other than PASS, ".", or empty are excluded. This removes variants flagged as low quality by the variant calling pipeline.
**Genotype validity:** Variants where the genotype (GT) field contains only missing alleles are excluded.
**Read depth:** If the DP FORMAT field is present, variants with depth below 10 are excluded to remove poorly supported calls.
**Genotype quality:** If the GQ FORMAT field is present, variants with genotype quality below 20 are excluded to remove uncertain genotype assignments.
Dosage for each passing variant is computed as the count of non-reference alleles: 0 (homozygous reference), 1 (heterozygous), or 2 (homozygous alternate). Both unphased (/) and phased (|) genotype encodings are handled. Chromosome identifiers are normalized by stripping the "chr" prefix if present and mapping "M" to "MT".
## PRS Scoring
### Raw Score
For each disease, the raw polygenic risk score is the weighted sum of per-variant dosage values:
$$
\text{PRS}_{\text{raw}} = \sum_{i=1}^{n} w_i \cdot d_i
$$
Where $w_i$ is the GWAS effect size (log odds ratio or beta coefficient) for variant $i$ and $d_i$ is the observed dosage (0, 1, or 2) from the VCF. When a variant in the disease weight set is not present in the input VCF, it is imputed at the expected dosage under Hardy-Weinberg equilibrium:
$$
d_i^{\text{imputed}} = 2 \cdot \text{EAF}_i
$$
Where $\text{EAF}_i$ is the effect allele frequency in the reference population. This imputation strategy is equivalent to assigning the population mean genotype for that variant and is standard practice in PRS computation with partially matched SNP sets.
Variant matching is attempted first by rsID, then by chromosomal position (GRCh38 chromosome:position) as a fallback. Per-disease statistics track how many variants were matched by rsID, matched by position, or imputed.
### Population Standardization
The raw PRS is standardized to a Z-score using theoretical population statistics derived from Hardy-Weinberg equilibrium. The population mean and standard deviation are computed analytically from the effect size and allele frequency information in the GWAS weight set:
$$
\mu = \sum_{i=1}^{n} 2 \cdot \text{EAF}_i \cdot w_i
$$
$$
\sigma = \sqrt{\sum_{i=1}^{n} 2 \cdot \text{EAF}_i \cdot (1 - \text{EAF}_i) \cdot w_i^2}
$$
The Z-score is then:
$$
Z = \frac{\text{PRS}_{\text{raw}} - \mu}{\sigma}
$$
This approach derives the reference distribution from the expected genetic variance under HWE, making the standardization independent of any empirical reference cohort for the raw computation step.
### Percentile and Risk Category
The Z-score is converted to a population percentile using the standard normal cumulative distribution function, calibrated against the UK Biobank European reference population (n = 488,000):
$$
\text{Percentile} = \Phi(Z) \times 100
$$
Percentiles are clamped to the range \[0.1, 99.9] to avoid boundary values. Individuals are then classified into one of four risk categories based on their percentile rank:
| Category | Percentile Range |
| --------- | ---------------- |
| Low | Below 20th |
| Average | 20th to 60th |
| High | 60th to 80th |
| Very High | 80th and above |
Relative risk estimates are retrieved from disease-specific quartile lookup tables derived from published GWAS findings. Each disease provides four relative risk values corresponding to the bottom 25%, 25th to 50th, 50th to 75th, and top 25th percentile bins of the score distribution.
## Disease Panels
RevRisk includes 36 built-in disease panels with curated GWAS lead SNPs across seven clinical categories. Each panel is linked to a PGS Catalog identifier for reference and includes disease-specific population prevalence and quartile-based relative risk values.
| Category | Diseases |
| --------------------------- | ------------------------------------------------------------------------------------------------------------------ |
| Cardiovascular | Coronary Artery Disease, Atrial Fibrillation, Heart Failure, Ischemic Stroke, Hypertension, Venous Thromboembolism |
| Neurological | Alzheimer's Disease, Age-Related Macular Degeneration, Parkinson's Disease, Multiple Sclerosis, Epilepsy, Migraine |
| Metabolic and Endocrine | Type 2 Diabetes, Type 1 Diabetes, Obesity, Hypothyroidism, PCOS |
| Oncology | Prostate Cancer, Breast Cancer, Colorectal Cancer, Lung Cancer, Pancreatic Cancer, Melanoma |
| Inflammatory and Autoimmune | Inflammatory Bowel Disease, Rheumatoid Arthritis, Systemic Lupus Erythematosus, Psoriasis, Celiac Disease |
| Respiratory and Renal | Asthma, COPD, Chronic Kidney Disease |
| Psychiatric | Schizophrenia, Bipolar Disorder, Major Depressive Disorder, ADHD |
Each SNP record in the built-in panels stores the rsID, GRCh38 chromosome and position, effect allele, other allele, effect size (log OR or beta), and effect allele frequency in the European ancestry reference population.
## PGS Catalog Integration
In addition to the 36 built-in panels, RevRisk integrates with the PGS Catalog REST and FTP APIs to support scoring against any of over 4,000 published polygenic scores. Users can search the catalog by trait keyword and select one or more external scores to include alongside the built-in panels.
When a custom PGS score is requested, the engine downloads the harmonized GRCh38 scoring file from the PGS Catalog FTP (falling back to the primary scoring file if the harmonized version is unavailable). Scoring files are cached locally for seven days so that repeated analyses of the same score do not require repeated downloads. The downloaded file is parsed to extract the same SNP-level fields used by the built-in panels, and the full PRS pipeline is applied identically. For custom scores where disease-specific relative risk quartile tables are unavailable, a conservative generic relative risk table is applied (Q1: 0.60, Q2: 0.85, Q3: 1.20, Q4: 2.00).
## Running the Engine
### Inputs
| Parameter | Required | Description |
| ---------- | -------- | -------------------------------------------------------------------------------------------------- |
| `vcf_file` | Yes | Per-sample genotype VCF (`.vcf`, `.vcf.gz`, `.bcf`) |
| `pgs_ids` | No | JSON array of custom PGS Catalog scores to include (e.g. `[{"pgs_id":"PGS000036","label":"T2D"}]`) |
When `pgs_ids` is omitted, all 36 built-in panels are scored using local weight tables with no network calls. When a PGS ID is provided that is not in the built-in set, the scoring file is downloaded from the PGS Catalog FTP and cached for 7 days.
### Outputs
Upon completion, the engine returns results for each disease scored:
* **Raw score:** Linear combination of effect sizes and genotype dosages
* **Z-score:** Standardized score relative to the HWE-derived population distribution
* **Percentile:** Population rank relative to UK Biobank reference, clamped to \[0.1, 99.9]
* **Risk category:** Low, Average, High, or Very High based on percentile thresholds
* **Relative risk:** Quartile-based relative risk estimate from disease-specific lookup tables
* **SNP match statistics:** Counts of variants matched by rsID, matched by position, and imputed per disease
**VCF QC statistics** are also returned: total variants, variants passing the FILTER field, variants excluded by insufficient depth, variants excluded by insufficient genotype quality, and variants excluded by invalid genotype calls.
**Alignment statistics** summarize matching across all diseases: total PRS SNPs, matched by rsID, matched by position, imputed, and overall match rate.
A summary object reports the number of diseases with elevated scores (above the 75th percentile), the number of diseases scoring in the Very High category (above the 80th percentile), the disease with the highest percentile, and the maximum percentile observed.
# RevScan
Source: https://docs.revilico.bio/docs/revscan
Optical Chemical Structure Recognition — Convert Chemical Structure Images to SMILES Notation
## Why Use RevScan?
RevScan converts images of chemical structures into machine-readable SMILES strings using optical structure recognition (OSRA). This is directly useful when extracting structures from publication figures, patent images, scanned laboratory notebooks, or any source where the chemical information exists as a raster image rather than a digital structure file. The resulting SMILES can be immediately used as input to any Revilico engine.
## Background
Chemical structure diagrams in the scientific literature are typically embedded as images in PDFs or scanned documents. Re-entering these structures manually into drawing tools is time-consuming and error-prone. Optical Structure Recognition (OSRA) is a computer vision approach that parses the geometry, bond angles, atom labels, and ring systems in a structure diagram image and reconstructs the corresponding connection table, producing a valid SMILES string.
RevScan uses OSRA (Optical Structure Recognition Application), an open-source tool developed at NCI/NIH, which is optimized for processing 2D chemical structure depictions from publications and patents.
## How It Works
The user uploads an image containing one or more 2D chemical structure drawings. RevScan passes the image through the OSRA recognition pipeline, which identifies atom positions, bond types (single, double, triple, aromatic), stereocenters, ring systems, and atom labels for heteroatoms and charges. The recognized connection table is then converted to a canonical SMILES string.
**Tips for best results:**
* Use clean, high-resolution images with clear bond lines and legible atom labels
* Images with dark structures on white backgrounds produce the most accurate recognition
* Avoid images with text overlapping structural elements
* Simple structures are recognized more accurately than complex polycyclic or organometallic structures
* Always verify the output SMILES against the source image, as recognition accuracy may not be 100% for complex or low-quality inputs
## Running the Engine
### Inputs
| Input | Description |
| ------------------------ | ------------------------------------------------------------------------- |
| Chemical structure image | JPEG, PNG, TIFF, or GIF file containing one or more 2D structure drawings |
Files can also be dragged and dropped directly from the Data Engineering environment.
### Outputs
* **SMILES string:** Canonical SMILES representation of the recognized chemical structure
* The output SMILES can be copied directly into any Revilico engine input field or saved to a review session
# RevScreen
Source: https://docs.revilico.bio/docs/revscreen
Optimizing Ligand Binding Predictions with Revilico's Docking Engines: Static, Flexible, and Ensemble Methods
## Why use this engine?
In this document we will use Revilico’s Virtual Screening Engine to dock a library of candidate ligands against a protein target in order to identify the most promising compounds for downstream analysis.. We begin with rigid docking to enable high-throughput screening and rapid downselection. Selected ligands are then evaluated using flexible docking,, which allows key protein residues to move and capture more realistic ligand-protein interactions.. Finally, ensemble docking is applied to assess ligand binding across multiple protein molecular dynamic conformations, providing insight into how ligands interact with the dynamic nature of the protein

## Background
Molecular docking is critical in chemistry for predicting the most stable, 3D orientation of a ligand within a protein binding sight along with it’s energetics and binding affinities, directly informing rational drug design, virtual screening, and understanding molecular drivers of binding. Running physical assays is a task that can be both expensive and very time consuming. Without computational tools such as docking, chemists would have to send incredibly large compound libraries into high throughput screening facilities in order to get a few hit compounds that effectively bind to the protein. Docking is an effective solution where chemists will be able to downselect their compound libraries to a much smaller batch of compounds with high predicted affinities, essentially cutting down on both time and cost to run experimental validation. In our own internal studies, we have observed almost a 20x improvement of our high throughput screening hit rates when utilizing these engines. We introduce Revilico’s Virtual Screening Engine, an engine composed of 3 different types of docking: Static Docking, Flexible Docking, and Ensemble Docking.
**Static Docking**
In static docking, we will predict ligand binding modes and relative affinities with a validated binding site using a fixed protein structure (e.g. the protein remains rigid whereas the ligand is flexible and can explore different poses). Static Docking is utilized for high throughput screens of larger libraries (>2M compounds). Overall, it is a GPU accelerated algorithm that allows for us to calculate binding scores and conformations of compounds at scale to better analyze vast chemical spaces without needing to run computationally extensive algorithms.
Before diving into how static docking works, we need to understand how we can calculate binding energy. It can be calculated with the following equation, composed of several summed energetic components at the atomistic level:
$$
E_{\text{binding}} = E_{\text{vdW}} + E_{\text{elec}} + E_{\text{hbond}} + E_{\text{desolv}} + E_{\text{tors}}
$$
$E_{\text{vdW}}$ is energy from Van der Waals calculated with the following equation
$$
E_{\text{vdW}} = \sum \left( \frac{A}{r^{12}} - \frac{B}{r^{6}} \right) \cdot S(r)
$$
Where A,B are Lennard-Jones parameters (atom-type specific), r is interatomic distance, and S(r) is a smoothing function.
$E_{\text{elec}}$ is the energy from electrostatics calculated with the following equation
$$
E_{\text{elec}} = \sum_{i
# RevSingleCell
Source: https://docs.revilico.bio/docs/revsinglecell
Prioritizing Therapeutic Candidates Using Revilico's scRNA Seq Engine: DEG Analysis, Automated Target ID, and Temporal Omics
## Why Use This Engine?
In the documentation below, we will use Revilico's Single Cell RNA sequencing analysis engine to explore and interpret a disease state from a transcriptomic perspective and evaluate whether a particular gene can serve as a potential therapeutic target.
## Background
Transcriptomics is the study of an organism's RNA molecules at a specific time in order to derive insights such as gene expression and gene regulation in many different scenarios ranging from cellular function to comprehension of a disease state. One type of transcriptomic technology in particular is Single Cell RNA sequencing, which is a powerful technique where gene expression of several transcripts can be measured in each individual cell across thousands of cells, allowing users to screen cells of different conditions at a high throughput level.
Revilico's scRNA-seq Analysis Engine integrates 3 different workflows where the user can do the following: (1) perform genome-wide differential expression screening (**DEG Analysis**), (2) evaluate therapeutic target candidates through gene regulatory network analysis, regulatory correlation mapping, and downstream cascade prediction (**Automated Target ID**), and (3) model dynamic gene expression transitions using pseudotime trajectory inference coupled with RNA velocity-based temporal modeling (**Temporal Omics Analysis**). We will explain these in greater detail below. This engine supports hypothesis generation, target validation, and mechanistic understanding of perturbation responses in biological systems.
Before diving into each workflow, we will first describe the core preprocessing pipeline that runs at the beginning of every analysis, as these steps are foundational to all three workflows described above.
## Core Preprocessing Pipeline
### 1. Data Loading and Format Support
The engine accepts several common single-cell data formats: `.h5ad` (native AnnData), `.h5` (10x Genomics HDF5 or generic HDF5), and count matrix files (`.csv`, `.tsv`, `.txt`). When performing differential expression between conditions, the user supplies both a control dataset and an experimental dataset. The engine concatenates these into a single AnnData object and assigns a `condition` label (control or experimental) to each cell in the observation metadata. This label is then used downstream during DEG analysis to group cells by biological condition rather than by cluster alone.
### 2. Quality Control
Once data is loaded, quality control filtering is applied to remove low-quality cells and uninformative genes. The engine computes the following per-cell metrics: number of detected genes, total UMI count, and the fraction of reads mapping to mitochondrial genes (genes whose names begin with "MT-"). Cells with fewer than `min_genes` or more than `max_genes` detected genes are removed, cells exceeding `max_pct_mt` percent mitochondrial content are removed, and genes detected in fewer than 3 cells are removed. The default thresholds are `min_genes = 200`, `max_genes = 6000`, and `max_pct_mt = 20%`. After filtering, QC summary metrics are reported including cells before and after filtering, median genes per cell, median UMI per cell, and median mitochondrial percentage. Optionally, doublet detection can be enabled via Scrublet, which flags likely multiplets before downstream analysis.
### 3. Normalization and Highly Variable Gene Selection
After QC, the raw count matrix is normalized so that each cell sums to 10,000 counts (counts per million scaling), then log-transformed:
$$
\tilde{x}_{ij} = \log\left(\frac{x_{ij}}{\sum_{j} x_{ij}} \cdot 10000 + 1\right)
$$
Where $x_{ij}$ is the raw count for gene $j$ in cell $i$ and $\tilde{x}_{ij}$ is the normalized log expression. The raw counts are preserved in `adata.raw` before any gene selection or scaling occurs. Highly variable genes (HVGs) are then selected using a mean-dispersion approach adapted from Seurat, retaining the top 2,000 genes with the highest normalized dispersion relative to their mean expression. Gene expression is then scaled to zero mean with a maximum value cap of 10 to prevent outlier genes from dominating variance decomposition.
Optionally, the engine can regress out technical covariates (total UMI count and mitochondrial percentage) before scaling, which is useful when these variables are expected to confound the biological signal of interest.
### 4. Dimensionality Reduction
Principal Component Analysis (PCA) is applied to the scaled HVG matrix using an ARPACK sparse solver, extracting up to 50 principal components. The number of components is automatically capped at min(n\_cells - 1, n\_genes). A k-nearest neighbor graph is then constructed in PCA space using 15 neighbors and 40 PCs, followed by UMAP embedding for 2D visualization.
If the user provides a batch key column in the cell metadata, Harmony batch correction is applied to the PCA embedding before neighbor graph construction. Harmony iteratively adjusts the PCA coordinates so that cells from different batches intermix based on biological identity rather than technical origin. The UMAP is recomputed on the Harmony-corrected representation, and both the original and corrected UMAPs are retained for comparison.
### 5. Clustering and Cell Type Annotation
The engine runs both Leiden and Louvain graph-based clustering algorithms on the neighbor graph. The resolution parameter (default 0.5) controls cluster granularity; higher values produce more fine-grained clusters. Leiden clustering is used as the primary cluster label throughout downstream analyses.
For each cluster, the top 50 marker genes are identified using a Wilcoxon rank-sum test against all remaining cells, ranked by test statistic. Cell type annotation is then performed by scoring the overlap between each cluster's marker genes and a curated database of canonical marker genes spanning immune cell types, stromal cells, epithelial cells, and neurons. The engine also handles gene identifiers in Ensembl format (ENSG, ENSMUSG) by converting them to gene symbols via the MyGeneInfo API or the Ensembl REST API before annotation. The cell type with the highest enrichment score against the marker database is assigned to each cluster.
## DEG Analysis
### Overview
The DEG Analysis workflow performs genome-wide differential expression testing between two biological conditions (e.g. treated vs. untreated, disease vs. healthy). The user uploads a control dataset and an experimental dataset, and the engine identifies which genes are significantly upregulated or downregulated in the experimental condition. The results include a ranked list of DEGs per group, a volcano plot, and a heatmap of the top differentially expressed genes.
### Wilcoxon Rank-Sum Test
By default, the engine uses the Wilcoxon rank-sum test (also called the Mann-Whitney U test) for differential expression. This is a non-parametric test that is well suited to the count-based, overdispersed distributions of scRNA-seq data. For each gene, cells in each group are ranked by expression, and the test statistic $U$ is computed as:
$$
U = R_1 - \frac{n_1(n_1 + 1)}{2}
$$
Where $R_1$ is the sum of ranks for cells in group 1 and $n_1$ is the number of cells in group 1. The normalized test statistic is then converted to a p-value. All p-values are corrected for multiple testing using the Benjamini-Hochberg false discovery rate procedure. As an alternative to the Wilcoxon test, the user can select a t-test for cases where parametric assumptions are acceptable.
Log fold-change between groups is computed on the log-normalized expression values:
$$
\text{logFC} = \overline{\log(x_{\text{exp}} + 1)} - \overline{\log(x_{\text{ctrl}} + 1)}
$$
Where $\overline{\log(x_{\text{exp}} + 1)}$ and $\overline{\log(x_{\text{ctrl}} + 1)}$ are the mean log-normalized expression values in the experimental and control groups, respectively. Genes with adjusted p-value below 0.05 are considered statistically significant. The engine reports separately the number of significantly upregulated genes (logFC > 0) and significantly downregulated genes (logFC \< 0).
### Grouping Strategy
When the user provides both a control and experimental dataset, DEG analysis is grouped by `condition` (control vs. experimental). If only a single dataset is provided, DEG analysis is performed per Leiden cluster against all other cells, which is equivalent to marker gene detection at the condition level.
### Outputs
* **Volcano plot:** Each gene is plotted with logFC on the x-axis and -log10(adjusted p-value) on the y-axis, colored by significance status
* **DEG heatmap:** Mean expression of the top differentially expressed genes per group, visualized as a clustered heatmap
* **DEG table:** Per-group ranked list of genes with score, adjusted p-value, and logFC
## Automated Target ID
### Overview
Automated Target ID builds on the DEG results to evaluate which differentially expressed genes are viable therapeutic targets. This workflow combines gene regulatory network analysis, regulatory correlation mapping, and downstream cascade prediction using transcription factor activity inference, pathway activity inference, and druggability annotation. The goal is to move from a ranked list of DEGs to a prioritized set of target candidates with mechanistic context.
### Transcription Factor Activity Inference (Regulatory Network Analysis)
Rather than relying on raw gene expression alone, the engine estimates transcription factor (TF) activity at the single-cell level using the Univariate Linear Model (ULM) method from the decoupler library. TF activity scores reflect how strongly a TF's known regulon is coordinately expressed in each cell, providing a more robust regulatory signal than the expression of the TF gene itself. The TF-target interaction network is sourced from Dorothea, a curated database of human TF-target interactions at confidence levels A through C.
For each TF $t$ and cell $c$, ULM fits the following linear model:
$$
y_c = a_t \cdot x_{t,c} + \varepsilon_c
$$
Where $y_c$ is the per-cell expression of all target genes in the regulon, $x_{t,c}$ is the prior weight (regulatory sign and confidence) of each target gene under TF $t$, $a_t$ is the estimated activity coefficient, and $\varepsilon_c$ is the residual. The activity score reported for each TF is the t-statistic of the coefficient $a_t$, which captures both the magnitude and consistency of regulon induction. Positive scores indicate that the TF's activating targets are coordinately upregulated; negative scores indicate that the TF's repressive program is engaged. Raw counts are used when available (from `adata.raw`) for this calculation to preserve the original signal.
### Pathway Activity Inference (Regulatory Correlation Mapping)
Pathway activity is estimated using the Multivariate Linear Model (MLM) from decoupler, applied to the PROGENy signaling pathway network. PROGENy is derived from a large compendium of perturbation experiments and provides a set of gene-level footprint weights for 14 core signaling pathways including EGFR, MAPK, PI3K, TGFb, TNF, WNT, and others. For each pathway $p$ and cell $c$:
$$
\mathbf{y}_c = \mathbf{X} \cdot \boldsymbol{\beta}_p + \boldsymbol{\varepsilon}_c
$$
Where $\mathbf{y}_c$ is the vector of expression values for pathway footprint genes in cell $c$, $\mathbf{X}$ contains the PROGENy weights for pathway $p$ (top 500 response genes), and $\boldsymbol{\beta}_p$ is estimated by ordinary least squares. The estimated coefficient $\beta_p$ serves as the per-cell pathway activity score. This approach links the observed transcriptional state of each cell to the activity of specific upstream signaling cascades, allowing the user to identify which pathways are most perturbed in the experimental condition and whether a target gene sits at the nexus of active regulatory networks.
### Downstream Cascade Prediction and Target Annotation
The top DEGs ranked by absolute logFC are queried against the MyGeneInfo API to retrieve structured biological annotations for each candidate target. For each gene, the engine retrieves the gene name and functional summary, associated KEGG pathways (top 3), Gene Ontology Biological Process terms, and the Pharos druggability tier (TDL). Pharos classifies targets into four tiers: Tclin (approved drug target), Tchem (has known small molecule ligands), Tbio (biologically characterized but not yet drugged), and Tdark (little known biology). This classification directly informs target tractability.
The combination of TF activity scores, pathway activity scores, and druggability annotation constitutes the regulatory context layer for each DEG. A gene that is differentially expressed, sits downstream of an active TF regulon, is linked to a perturbed signaling pathway, and has a favorable Pharos tier represents a well-supported therapeutic candidate.
### Outputs
* **TF activity heatmap:** Per-cell TF activity scores across inferred TFs, shown as a heatmap
* **TF activity UMAP:** UMAP embeddings colored by individual TF activity scores
* **Pathway activity heatmap:** Per-cell pathway activity scores across the 14 PROGENy pathways
* **Pathway activity UMAP:** UMAP colored by individual pathway activity
* **Target table:** Annotated list of top candidate targets with KEGG pathways, GO terms, and druggability tier
## Temporal Omics Analysis
### Overview
The Temporal Omics Analysis workflow models how gene expression states transition over time, reconstructing the dynamic trajectory of cells as they progress through a biological process such as differentiation, treatment response, or disease progression. The engine infers a pseudotime ordering of cells and identifies which genes are most strongly coupled to that temporal axis. Two methods are available depending on the data: RNA velocity-based pseudotime using scVelo when spliced and unspliced RNA counts are available, and Diffusion Pseudotime (DPT) as a robust fallback for standard count matrices.
### scVelo RNA Velocity (ODE-Based Temporal Modeling)
When the input data contains separate layers for spliced and unspliced RNA counts (as produced by tools such as velocyto or STARsolo), the engine applies scVelo's dynamical model. This model treats transcription, splicing, and degradation as a coupled ODE system, recovering the kinetic parameters that best explain the observed spliced-to-unspliced ratios across cells.
The governing equations for a given gene are:
$$
\frac{du}{dt} = \alpha(t) - \beta \cdot u
$$
$$
\frac{ds}{dt} = \beta \cdot u - \gamma \cdot s
$$
Where $u$ is the unspliced (pre-mRNA) count, $s$ is the spliced (mature mRNA) count, $\alpha(t)$ is the transcription rate (which switches between an active state $\alpha > 0$ and a quiescent state $\alpha = 0$), $\beta$ is the splicing rate constant, and $\gamma$ is the mRNA degradation rate constant. By fitting these parameters from the data, scVelo identifies each gene's current transcriptional state (inducing or repressing) and projects a velocity vector for each cell in PCA space, indicating the direction of future gene expression change.
RNA velocity vectors are projected onto the UMAP embedding to provide an intuitive directional visualization of cell state transitions. Latent time is then inferred from the reconstructed dynamics, assigning each cell a continuous pseudotime value between 0 and 1 that reflects its position along the inferred transcriptional trajectory. The root cluster is automatically assigned as the cluster with the lowest mean latent time.
### Diffusion Pseudotime (DPT)
For datasets without spliced or unspliced layers, the engine falls back to Diffusion Pseudotime. DPT constructs a diffusion operator over the cell-cell similarity graph by normalizing the affinity matrix:
$$
M = D^{-1/2} K D^{-1/2}
$$
Where $K$ is the Gaussian kernel affinity matrix computed from the PCA representation and $D$ is the diagonal degree matrix with $D_{ii} = \sum_j K_{ij}$. The eigendecomposition of $M$ yields diffusion components that capture the global geometry of the data manifold. Pseudotime is then defined along the principal diffusion axis, anchored to a root cell selected from the root cluster as the cell with the highest value in the first diffusion component. Each cell receives a pseudotime value in \[0, 1] proportional to its diffusion distance from the root.
### Lineage Gene Identification
After pseudotime is computed by either method, the engine identifies genes whose expression is significantly correlated with pseudotime within each cluster. For each cluster, Pearson correlation is computed between each gene's expression and pseudotime across cells in that cluster. Genes with absolute Pearson $r > 0.2$ and significant p-values are reported as lineage-associated genes for that cluster, providing a ranked list of genes that drive or mark the transition along the inferred trajectory.
### Outputs
* **Pseudotime UMAP:** UMAP colored continuously from early (0) to late (1) pseudotime
* **Trajectory plot:** UMAP with RNA velocity stream arrows (scVelo) or pseudotime contours (DPT), with cluster ordering by mean pseudotime
* **Cluster pseudotime table:** Mean pseudotime per cluster, reflecting temporal ordering of cell populations
* **Lineage gene table:** Per-cluster list of genes most correlated with pseudotime, with Pearson $r$ and p-value
## Running the Engine
### Inputs
To configure and launch an analysis, the user provides the following:
| Parameter | Default | Description |
| ---------------------- | -------- | ------------------------------------------------------ |
| Control dataset | Required | `.h5ad`, `.h5`, `.csv`, `.tsv`, or `.txt` count matrix |
| Experimental dataset | Optional | Second dataset for condition-level DEG analysis |
| `min_genes` | 200 | Minimum genes per cell for QC filtering |
| `max_genes` | 6000 | Maximum genes per cell for QC filtering |
| `max_pct_mt` | 20.0 | Maximum mitochondrial gene percentage per cell |
| `resolution` | 0.5 | Leiden clustering resolution |
| `batch_key` | None | Metadata column for Harmony batch correction |
| `run_harmony` | False | Enable Harmony batch correction |
| `regress_out` | False | Regress out total counts and mitochondrial percentage |
| `run_tf_activity` | False | Enable TF activity inference (ULM + Dorothea) |
| `run_pathway_activity` | False | Enable pathway activity inference (MLM + PROGENy) |
| `run_dge` | False | Enable DEG analysis |
| `run_target_id` | False | Enable automated target ID (requires DEG analysis) |
| `run_temporal_omics` | False | Enable trajectory and pseudotime analysis |
| `dge_method` | wilcoxon | DEG test method (`wilcoxon` or `t-test`) |
| `dge_top_n` | 50 | Number of top DEGs to report per group |
### Outputs
Upon completion, the engine reports the following summary statistics and result objects:
* **QC metrics:** Cells before and after filtering, median genes, median UMI, median mitochondrial percentage
* **Cluster summary:** Number of Leiden clusters, cells per cluster, and assigned cell type per cluster
* **Marker genes:** Top 50 Wilcoxon-ranked marker genes per cluster with logFC and adjusted p-value
* **Cell type composition:** Cell type assignments with counts and proportions
* **DEG results:** Per-group ranked gene lists with test statistic, logFC, and adjusted p-value (if enabled)
* **TF activity results:** Per-cell activity scores for all inferred TFs (if enabled)
* **Pathway activity results:** Per-cell activity scores for all 14 PROGENy pathways (if enabled)
* **Target annotation table:** Top candidate genes with KEGG pathways, GO terms, and Pharos druggability tier (if enabled)
* **Trajectory results:** Pseudotime per cell, cluster ordering, inference method, and lineage-correlated genes (if enabled)
* **Plots:** UMAP by cluster and cell type, QC violin plots, marker gene heatmap, harmony comparison, TF heatmaps, pathway heatmaps, volcano plot, DEG heatmap, pseudotime UMAP, trajectory plot
### AI Copilot
Results are accessible through an interactive AI copilot powered by Claude. The copilot has direct access to all analysis results and can generate on-demand visualizations, retrieve marker gene lists, compare cell type compositions, query differential expression between specific clusters, and summarize trajectory findings. The copilot uses tool-use to call plot generation and data retrieval functions in real time, so users can explore their data through natural language queries without needing to rerun the pipeline.
# RevSol
Source: https://docs.revilico.bio/docs/revsol
A Crash Course on Revilico’s Compound Solubility Engine
## Why Use this product?
Solubility is a core property that should be computationally tested in a way that represents biological and lab based conditions for downstream experimentation. In later lead optimization stages, many compounds may fail due to bad solubility in certain solvent conditions which has potential to cause downstream issues with bioavailability, assay reliability, and formulations. Revilico’s Solubility Engine allows chemists to test a compound's intrinsic solubility across temperature gradients and different solvents, quickly and effectively.

## Background
Compound solubility is an important principle for evaluating the effectiveness of the delivery of a small molecule, as the molecule must be sufficiently dissolved to be absorbed, distributed and tested reliably. Poor solubility can limit bioavailability, cause assay failures, complicate formulation, and ultimately prevent an otherwise potent compound from becoming a viable drug. Revilico’s Compound Solubility Engine proves to be an effective tool for a high throughput screening of several molecules particularly measuring for solubility, saving both time and money for downstream experimentation.
Diving down into the mathematical theory behind solubility prediction, the binding free energy of solvation as a function of temperature can be defined by the following equation:
$$
\Delta G(T) = \Delta H - T \Delta S
$$
Where (Delta H) is the enthalpic term which is the net energy change resulting from breaking and forming of chemical interactions (i.e. bonds, intermolecular forces) and (Delta S) is the entropic term which measures the shift in molecular disorder, randomness, or energy dispersal during a physical or chemical process. In order for us to calculate this property across several compounds, it can become computationally intensive. This would require us to run molecular dynamics and free energy calculation engines at scale, and due to high computational cost, it is infeasible for speedy on-demand screening. However, alternative deep learning models can be utilized to calculate these properties at scale across several different conditions.
Now diving into the workflow behind this engine, computing solubility from first principles will be a very difficult task where we will need to model the solid crystal lattice, model the solvent explicitly, simulate molecules leaving the crystal, and sample many configurations, a task that would be very computationally expensive. Instead we will measure solubility using a learned thermodynamic model that is trained on experimental solubility data. Essentially the mapping of this model to predicted solubility free energy can be defined by the following equation:
$$
(\text{molecule}, \text{solvent}, T) \rightarrow \Delta G_{\text{solv}}(T)
$$
This model would take our molecule and extract out molecular size and shape, polarity, hydrogen bond donor/acceptors, aromaticity, flexibility, approximate charge distribution, and whether it is a hydrophobic vs hydrophilic surface through a series of “features” that will be fed into a neural network. Utilizing this data and the molecular featurizations we create, we aim to understand the general ability for the compound to dissolve in its subsequent solvents.
The model will then take the solvent and extract out polarity, dielectric constant, hydrogen bonding ability, and cohesive energy to ensure molecular features that are extracted will match well with certain solvent features.
After encoding both our molecule and solvent, we will then encode temperature, apply a trained neural network to predict our (Delta G) of solvation, then convert this (Delta G) to solubility as a result of the learned or calibrated relationships based on the data.
## Interactive Results Viewer
Explore compound solubility predictions interactively. Select a molecule to view its predicted solubility across different solvents and temperatures.
# RevTrain
Source: https://docs.revilico.bio/docs/revtrain
Model Training for Generative Chemistry Campaign with Revilico’s Custom Model Training Engine
## Why Use this product?
Revilico’s Custom Model Training Engine is a transfer learning tool that applies pre-trained models to your curated compound library creating custom models that generate chemistry aligned with your organization's structural preferences. IP landscape, and SAR knowledge for downstream library generation and lead optimization campaigns. This engine is best used when you need to fine-tune the De Novo Library Generation Engine or the Molecular Optimization engine, enabling the generation of molecules that statistically resemble your training dataset rather than generic drug-like compounds.This feature helps you to establish your ‘chemical prior’ model that can be used for molecular optimization, allowing you to sample from the chemical space of your input molecular library and surrounding spaces.

## Background
When running the De Novo Library Generation Engine and the Molecular Optimization engine, the basis for generation of our datasets is the prior model which encapsulates a large chemical space. However, if our task requires us to only generate molecules within a certain chemical space. We will not be able to use the reinvent prior model, and we will now need a model that is only trained on that specific chemical space we are interested in. This is where Revilico’s Custom Model Training Engine comes into place. The goal of this engine is to take a compound library and transform it into a model that can be used downstream in De Novo Library Generation or in Molecular Optimization.
How does it work? We have two types of library generators. With the De Novo Generation generator, what we will do is input a compound library, and the output will be molecules that are similar in that chemical space. With Molecular Optimization Transformer, we will input a compound library, typically one that has a core scaffold with slight differences or medicinal chemistry transformations. This model will look at the transformation of the molecules within the library, focused on generating chemically plausible analogs that reflect the learned transformation patterns within that lead series or compound set.
With both of these models, it generally operates by first starting with our start token or some fragment of our SMILES string. The job of the model is to predict the next token given the previous tokens. It can be denoted by this probability.
$$
P(\text{next token} \mid \text{previous token}, \text{context})
$$
With De Novo Generation it can be more precisely defined as:
$$
P(\text{next token} \mid \text{previous token})
$$
Where the statistics come from all molecules in the training set. With Molecular Optimization it can be defined as:
$$
P(\text{output token} \mid \text{input molecule tokens}, \text{previous output tokens})
$$
Which means it is conditioned on the specific input molecule, answering the question, given this molecule, how do chemists usually modify it in this project, through learned associations of the chemical space.
During training the model is corrected in this process where the model’s prediction is compared to the actual next token in your dataset. If it guesses wrong, its internal weights are adjusted slightly. This will use cross-entropy loss as the loss function where the goal of the model is to minimize the cross entropy loss. Cross entropy loss can be denoted by the following formula:
$$
L_{CE} = - \sum_i y_i \log(p_i)
$$
Where i is the index token y is the true label as denoted by 1 or 0 and p is the predicted probability for token i.
We will end up with a model with updated model weights in accordance to the input training data. This model will now generate new models from scratch with either goal in mind of creating molecules that resemble a particular set or analog molecules similar to our training data set.
# RevTS
Source: https://docs.revilico.bio/docs/revts
DFT Transition State Search, Optimization, and Reaction Barrier Characterization
## Why Use This Engine?
In the documentation below, we will use Revilico's RevTS engine to locate and characterize transition states for chemical reactions involving small molecules or reactions occurring within protein binding sites. RevTS predicts the activation barrier (DG double-dagger) and reaction energy (DG) using a multi-stage pipeline that combines fast semiempirical nudged elastic band calculations for an initial pathway guess with DFT saddle-point optimization, frequency analysis, and intrinsic reaction coordinate (IRC) following for full characterization at the wB97M-V/def2-TZVP level. The resulting energetics are directly comparable to RevQS binding scores, enabling a unified quantum mechanical picture of both ground-state binding and reactive transformation within the same target.
## Background
A transition state (TS) is the highest-energy point on the minimum energy path connecting reactants to products. It corresponds to a first-order saddle point on the potential energy surface: a maximum along the reaction coordinate and a minimum in all other directions. Characterizing the TS provides the activation barrier DG double-dagger, which determines the rate of the reaction through the Arrhenius and Eyring equations, and the reaction energy DG, which determines thermodynamic favorability.
In drug discovery, transition state analysis is relevant to understanding metabolic reactions, covalent inhibitor mechanisms, enzymatic reaction barriers, and the kinetics of drug-induced conformational changes. RevTS uses xTB for fast semiempirical path generation, Psi4 for high-quality saddle-point optimization and IRC, and PySCF (the same engine as RevQS) for the final energy evaluation at wB97M-V/def2-TZVP, ensuring that binding energies and reaction barriers are computed at a consistent level of theory.
## Pipeline
**Stage 1: NEB Pathway Generation**
A nudged elastic band (NEB) calculation at the GFN2-xTB semiempirical level generates an initial estimate of the reaction path and transition state geometry. NEB constructs a chain of 8 molecular images connecting the reactant and product geometries, connected by spring forces that maintain equal spacing along the path while allowing each image to relax toward the minimum energy path. The spring potential is:
$$
V_{\text{spring}} = \frac{k_{\text{spring}}}{2}(|\mathbf{R}_{i+1} - \mathbf{R}_i| - |\mathbf{R}_i - \mathbf{R}_{i-1}|)^2
$$
With a default spring constant of 0.02 Eh/Bohr-squared. The highest-energy image along the NEB path serves as the initial TS guess for the DFT optimization. If xTB is unavailable, a linear interpolation of the Cartesian coordinates between reactant and product is used as the fallback starting path.
**Stage 2: DFT Saddle-Point Optimization**
The xTB TS guess is refined to a true first-order saddle point using the P-RFO (Partitioned Rational Function Optimization) algorithm in Psi4. P-RFO partitions the Hessian eigenvectors into a reaction mode (maximized) and all orthogonal modes (minimized), ensuring convergence to the correct saddle point rather than drifting toward a higher-order critical point or a minimum. The optimization uses wB97M-D3BJ/def2-SVP (Psi4 uses the D3BJ empirical dispersion variant of wB97M because Psi4 does not natively implement VV10 non-local correlation). Geometric differences between the D3BJ and full VV10 variants are below 0.3 degrees in angles and 0.005 Angstroms in bond lengths for typical organic reactions, making this combination appropriate for geometry optimization.
In binding-site mode, protein atoms outside the reactive region are frozen, allowing only the ligand and the immediately surrounding residue atoms to relax during the optimization.
**Stage 3: Frequency Analysis**
At the converged saddle point geometry, analytic second derivatives are computed to obtain the harmonic vibrational frequencies. A valid transition state has exactly one imaginary frequency (reported in cm-1 as a negative number by convention), corresponding to the normal mode that connects reactants to products along the reaction coordinate. Frequencies are scaled by a factor of 0.97 to correct for systematic overestimation of harmonic force constants. The zero-point energy correction to the barrier is extracted from the real frequencies.
**Stage 4: IRC Following**
The intrinsic reaction coordinate (IRC) confirms that the TS connects the intended reactants and products. Starting from the TS geometry and following the mass-weighted gradient in both directions using the Gonzalez-Schlegel second-order (GS2) method:
$$
\mathbf{R}(s + \delta s) = \mathbf{R}(s) - \delta s \frac{\mathbf{g}(s)}{|\mathbf{g}(s)|}
$$
Where the step is taken in mass-weighted Cartesian coordinates with step size 0.1 Bohr times amu to the one-half. The IRC terminates when the gradient falls below the convergence threshold, confirming that each direction leads to a proper minimum. If Psi4 is unavailable, a first-order Euler predictor IRC in PySCF provides the same confirmation at reduced accuracy.
**Stage 5: High-Level Single-Point Energy**
The final energetics are computed at the wB97M-V/def2-TZVP level using PySCF with GPU4PySCF acceleration, the same method as RevQS. For binding-site calculations, the Boys-Bernardi counterpoise correction is applied. The activation barrier and reaction energy are then:
$$
\Delta G^{\ddagger} = E(\text{TS, wB97M-V/TZVP}) - E(\text{reactant, wB97M-V/TZVP}) + \Delta E_{\text{ZPE}} + \Delta E_{\text{CP}}
$$
$$
\Delta G = E(\text{product, wB97M-V/TZVP}) - E(\text{reactant, wB97M-V/TZVP})
$$
**Wigner Tunneling Correction**
For reactions involving hydrogen transfer, quantum tunneling can significantly enhance the reaction rate beyond the classical Arrhenius contribution. The Wigner correction factor is:
$$
\kappa_W = 1 + \frac{1}{24}\left(\frac{h\nu^{\ddagger}}{k_BT}\right)^2
$$
Where $\nu^{\ddagger}$ is the magnitude of the imaginary frequency, $h$ is Planck's constant, and $k_B$ is the Boltzmann constant at 298.15 K. The effective rate constant is multiplied by $\kappa_W$.
## Running the Engine
### Inputs
| Parameter | Default | Description |
| ----------------- | ----------------- | ------------------------------------------------ |
| Reactant geometry | Required | XYZ or SDF file with 3D coordinates |
| Product geometry | Required | XYZ or SDF file with 3D coordinates |
| Mode | small\_molecule | `small_molecule` or `binding_site` |
| Protein PDB | Binding-site only | Protein structure for pocket context |
| Ligand SDF | Binding-site only | Ligand file for CP correction |
| NEB images | 8 | Number of path images for NEB |
| NEB method | gfn2 | Semiempirical level for NEB |
| TS method | psi4 | DFT engine for OptTS and IRC (`psi4` or `pyscf`) |
| Screen basis | def2-SVP | Basis for OptTS and IRC |
| High-level basis | def2-TZVP | Basis for final single-point energies |
| IRC step size | 0.1 | Mass-weighted IRC step in Bohr/amu-half |
| Pocket cutoff | 5.0 A | Residue radius for binding-site mode |
| Charge | 0 | Total molecular charge |
| Multiplicity | 1 | Spin multiplicity |
### Outputs
* **DG double-dagger:** Activation barrier in kcal/mol with ZPE and CP corrections
* **DG reaction:** Reaction energy in kcal/mol
* **Wigner kappa:** Tunneling correction factor
* **TS confirmation:** Whether a single imaginary frequency was found and IRC connected endpoints
* **Imaginary frequency:** Transition state normal mode frequency in cm-1
* **IRC profile:** Energy vs. reaction coordinate plot from TS to reactant and product
* **Per-residue decomposition:** Interaction contributions in binding-site mode
* **TS geometry file:** Optimized transition state structure in XYZ format
* **Method string:** Full description of the computational protocol applied
# RevViability
Source: https://docs.revilico.bio/docs/revviability
In Silico Drug Sensitivity and Cell Viability Prediction Across Cancer Cell Lines
## Why Use This Engine?
In the documentation below, we will use Revilico's RevViability engine to predict how a drug will affect cancer cell viability at different concentrations without running a wet-lab assay. RevViability generates full dose-response curves, IC50 values, sensitivity classifications, and selectivity indices across a panel of cancer cell lines, enabling researchers to prioritize compounds and cell line selections before committing to experimental resources.
## Background
Cell viability assays measure the fraction of living cells after drug treatment at a range of concentrations, producing a dose-response curve that is fit to the Hill sigmoid model to extract the IC50 (the concentration at which 50% of cells are killed). Running these assays experimentally across large compound libraries and diverse cell line panels is time-consuming and costly. The Genomics of Drug Sensitivity in Cancer (GDSC) project has generated over 467,000 drug-cell line viability measurements across 408 drugs and 978 cell lines, providing a large-scale training resource for predictive models.
RevViability uses a GNN+MLP hybrid model trained on GDSC1, GDSC2, and DepMap CCLE 24Q4 data. The graph neural network encodes the molecular structure of the drug as a graph of atoms and bonds, learning structural features relevant to cytotoxicity. The MLP combines these drug features with genomic features of the cell line (gene expression, copy number, mutation status) to predict the dose-response relationship. The model achieves R-squared of 0.87 on held-out test data against experimental GDSC values. Predictions are available for 972 CCLE-mapped cell lines.
## Dose-Response Modeling
The predicted dose-response relationship for each drug-cell line pair follows the Hill equation:
$$
v(c) = 1 - \frac{c^n}{c^n + \text{IC}_{50}^n}
$$
Where $v(c)$ is the viability fraction at concentration $c$, $n$ is the Hill coefficient controlling curve steepness, and IC50 is the half-maximal inhibitory concentration. The curve is evaluated at 64 concentration points spanning the user-defined range.
**Exposure time correction:** The base model IC50 corresponds to the GDSC standard 72-hour exposure. When the user selects a different exposure time $t$, the IC50 is corrected by:
$$
\text{IC}_{50}^{\text{eff}} = \text{IC}_{50}^{\text{base}} \times \left(\frac{72}{t}\right)^{0.4}
$$
This correction reflects the empirical relationship between exposure duration and apparent potency.
**Biological variability:** When replicates are requested, Gaussian noise calibrated to match GDSC inter-replicate coefficient of variation is added to the viability predictions, and IC50 confidence intervals are computed from the replicate distribution.
## Sensitivity Classification
IC50 values are classified into four sensitivity categories based on clinically relevant concentration thresholds:
| IC50 | Category | Interpretation |
| ------------ | ---------------- | ---------------------------------------------------------- |
| Below 0.1 uM | Highly Sensitive | Clinically achievable at low dose; likely on-target effect |
| 0.1 to 1 uM | Sensitive | Achievable plasma concentration for most drugs |
| 1 to 10 uM | Intermediate | May require high dose; off-target effects possible |
| Above 10 uM | Resistant | Unlikely to be clinically useful for this cell line |
## Selectivity Index
When multiple cell lines are profiled, the selectivity index quantifies how selective a drug is for the most sensitive cell line relative to the least sensitive:
$$
\text{SI} = \frac{\text{IC}_{50}^{\max}}{\text{IC}_{50}^{\min}}
$$
A high SI indicates on-target selectivity (one lineage is significantly more sensitive than others), which is the expected signature of a mechanism-based inhibitor. A low SI indicates pan-cytotoxicity.
## Media pH Simulation
RevViability models media acidification over the exposure time based on CO2 production and lactate accumulation from cellular metabolism. If pH drops below 6.8, a viability penalty factor is applied to correct for acidic stress-induced cytotoxicity that is independent of the drug mechanism. This provides more accurate predictions for long-duration assays at high cell densities.
## GDSC Reference Comparison
For drugs in the GDSC training database, the predicted IC50 is compared against the published experimental reference value and the fold deviation is reported. A fold error below 2x indicates excellent agreement. Fold errors between 2x and 5x are within the range of expected inter-laboratory variability in experimental viability assays.
## Batch Screening
RevViability supports screening up to 50 compounds simultaneously against any selection of the 972 available cell lines. Batch inputs accept compound name plus SMILES (one per line) or direct SMILES input. Results are returned as a sortable table of IC50 values and sensitivity classifications per compound per cell line, with CSV download for downstream analysis.
## Plate Layout and Assay Configuration
The engine simulates a standard 96-well plate format. Users assign cell lines and concentrations to individual wells or groups of wells using serial dilution presets, row and column selectors, or manual address entry. Four-cell-line layouts automatically fill all 96 wells with a serial dilution for each of four cell lines across two rows each. Well-level annotations can be assigned for control labeling and audit trail purposes.
## Running the Engine
### Inputs
| Parameter | Default | Description |
| ---------------------- | ----------------------------- | ----------------------------------------------------- |
| Drug SMILES | Required | Structure of the compound to test |
| Cell lines | A549, MCF7, MDA-MB-231, HEPG2 | Up to 972 CCLE-mapped cancer cell lines |
| Exposure time | 72 h | Duration of drug treatment (6 to 120 hours) |
| Hill coefficient | 1.2 | Dose-response curve steepness |
| Biological variability | None | `none`, `low`, `medium`, or `high` noise level |
| Replicates | 3 | Number of replicate curves (activates CI computation) |
| Concentration range | 0.001 to 100 uM | Dose range for the 64-point curve |
| Seeding density | 10,000 cells/well | Affects pH simulation and absolute viability scale |
### Outputs
* **IC50 values:** Exposure-corrected IC50 per cell line with 95% confidence interval
* **Dose-response curves:** 64-point Hill equation fits per cell line
* **Sensitivity classification:** Highly Sensitive, Sensitive, Intermediate, or Resistant badge per cell line
* **Selectivity index:** Ratio of maximum to minimum IC50 across all profiled cell lines
* **AUC:** Area under the dose-response curve (lower = more sensitive overall)
* **GDSC comparison:** Fold deviation from experimental reference IC50 where available
* **Waterfall plot:** Cell lines rank-ordered by IC50 with error bars
* **Heatmap:** Full 96-well plate colored by predicted viability
* **Lineage plot:** IC50 by tissue type showing tissue-level selectivity
* **ln(IC50):** Log-scale IC50 matching GDSC ln\_ic50 column format for direct database comparison
* **Protocol export:** Plain-text lab notebook format with plate layout, parameters, and results
* **CSV export:** Full well-level data with concentrations, viability values, IC50, and annotations
# Beginning Your First End-to-End Campaign
Source: https://docs.revilico.bio/guides/beginning-your-first-end-to-end-campaign
A step-by-step guide to standing up your first full computational + experimental drug discovery campaign with Revilico — from therapeutic strategy through parallel execution.
## Overview
Your first campaign with Revilico is a joint effort: our team works alongside you to lock in strategy, mine prior knowledge, and scope assays in parallel with the computational build-out, so that by the time compounds are ready for down-selection, everything downstream — CRO quotes, protein production, assay validation — is already lined up. This guide walks through the six steps we take together to get a campaign fully situated, from first conversation to parallel execution.
| Step | Name | Owner |
| ---- | ----------------------------------------------- | --------------------------- |
| 1 | Therapeutic Strategy & Positioning | Joint |
| 2 | Prior Knowledge Calibration | Revilico |
| 3 | Assay Strategy (First-Line & Early Second-Line) | Joint, in parallel with CRO |
| 4 | Computational Campaign Design | Revilico |
| 5 | Project Setup & Centralization | Revilico |
| 6 | Parallel Execution | Joint |
Steps 3 and 4 run **in parallel**, not sequentially. The goal is that the only thing standing between strategy and a confirmed lead set is a single med-chem down-selection step at the end.
***
## Step 1 — Finalize Therapeutic Strategy & Positioning
Before any computational or experimental work begins, we align on the shape of the campaign itself.
* **Therapeutic strategy** — the disease area, mechanism of action, and modulation approach (inhibition, activation, allosteric modulation, PPI disruption, etc.) you intend to pursue.
* **Market positioning** — where this program sits relative to existing and pipeline therapies, and what differentiates it.
* **Potential ROI** — a working estimate of commercial opportunity that justifies the scope of computational and experimental investment.
* **Pocket of interest** — the specific binding site, domain, or interface you want to drug, and why it's tractable.
This step produces the therapeutic hypothesis that every subsequent step is built around.
***
## Step 2 — Identify All Prior Knowledge
Before designing a single assay or docking run, we mine everything already known about your target so the campaign is calibrated against real biology and chemistry from day one — the same process we ran for your TNFα program, including the full prior compound set pull.
All available co-crystal structures for your target, including the bound ligand identity and resolution. These define validated binding poses and active-site conformations.
Residues known from structural or mutagenesis data to be mechanistically critical — contact residues from co-crystals, alanine-scan hot spots, or literature-annotated catalytic residues.
Compounds with documented activity against your target — pulled from ChEMBL, PubChem BioAssay, patents, and the literature.
For each known binder: the reported activity value (IC₅₀, Kᵢ, Kd, EC₅₀) and exactly how it was measured — assay format, biochemical vs. cellular, and conditions.
This prior-knowledge set becomes the calibration foundation referenced throughout Steps 3 and 4 — it's what lets us benchmark both computational engines and assay formats against ground truth before committing to production scale.
***
## Step 3 — Identify First-Line & Early Second-Line Assays
This is the first thing that needs to be locked down, and it runs **in parallel with the computational build-out** in Step 4 — so that the only remaining decision downstream is a med-chem-reviewed down-selection of compounds.
We finalize the assay strategy together with your CRO. For each assay under consideration, we work through:
What are we actually trying to measure — binding affinity, functional inhibition/activation, cellular potency, selectivity? Define the readout before choosing the format.
What reagents does the assay require, and are they already available, or do they need to be sourced or synthesized?
Which protein constructs need to be expressed? From-scratch expression or off-the-shelf/commercial protein — and if from scratch, what expression system and timeline?
What tags are required for the assay format (e.g. C-terminal 6×His for SPR), and which protein variants are needed (wild-type plus any relevant mutants)?
Single-dose screening for rapid triage, or full dose-response for confirmed hits? This determines throughput and cost per assay round.
Is there a validated kit or off-the-shelf assay we can use directly, or does the assay need to be set up and validated internally before it can run at scale?
Running this step in parallel with computational design (Step 4) means CRO lead times, protein expression, and assay validation are already progressing while the compound funnel is being built — so nothing sits idle waiting on the other.
***
## Step 4 — Define the Computational Campaign
We help finalize the computational strategy with you, including:
* **Docking configuration** — docking box definition and coordinates, and the exhaustiveness setting for each docking stage.
* **Generative engines** — which generative chemistry engines are needed to expand or design around your initial chemical matter.
* **MD analyses** — which molecular dynamics analyses are required (protein-in-water conformational sampling, protein-ligand stability, membrane permeability, etc.).
* **FEP** — whether free energy perturbation is warranted for this program, and at what stage (typically post hit-to-lead, on a narrowed set).
Alongside the strategy itself, we provide recommended next steps and forward you video tutorials and written guides so your team can execute directly in the Revilico platform as the campaign progresses.
***
## Step 5 — Set Up Your Project
Once Steps 1–4 are scoped, we create a dedicated **Revilico OS Project** for the target. This project becomes the single place where everything lives:
* The full computational campaign — docking runs, MD, FEP, generative design
* The full experimental campaign — assay results, CRO deliverables
* The therapeutic strategy and plan documented in Steps 1–2
* All CRO quotes, which we can help design on your behalf
Centralizing the campaign here means every stakeholder — computational, experimental, and CRO-facing — is working from the same source of truth as the program progresses.
***
## Step 6 — Move Into Parallel Execution
With strategy locked, prior knowledge mined, assays scoped, computational design finalized, and the project centralized, we move into execution — computational and experimental workstreams running **simultaneously** rather than sequentially.
Because the assay strategy (Step 3) and computational design (Step 4) were finalized in parallel rather than in sequence, this next iteration moves significantly faster than a first-time campaign: the only gate between the computational funnel and wet-lab confirmation is a single medicinal chemistry review and down-selection of the recommended compound set.
***
## Quick Reference Checklist
**Step 1 — Therapeutic Strategy**
* Define therapeutic strategy, mechanism, and modulation approach
* Establish market positioning and potential ROI
* Identify the pocket of interest
**Step 2 — Prior Knowledge**
* Pull all available co-crystal structures
* Identify key / hot-spot residues
* Compile known binders with activity values and assay provenance
**Step 3 — Assay Strategy (parallel with Step 4)**
* Define key measurements per assay
* Confirm required reagents and protein constructs
* Decide from-scratch expression vs. off-the-shelf protein
* Confirm tags and variants needed
* Decide single-dose vs. dose-response format
* Confirm whether assay validation is needed or a kit is available
* Finalize with CRO
**Step 4 — Computational Campaign**
* Define docking boxes, coordinates, and exhaustiveness
* Identify generative engines needed
* Scope required MD analyses
* Determine if/when FEP is warranted
**Step 5 — Project Setup**
* Create dedicated Revilico OS Project for the target
* Centralize computational + experimental plans, strategy, and CRO quotes
**Step 6 — Execution**
* Launch computational and experimental workstreams in parallel
* Reserve med-chem review and down-selection as the final gating step
# Filtering Molecules for your Virtual Screen
Source: https://docs.revilico.bio/guides/filtering-molecules-virtual-screen
A complete, engine-by-engine protocol for filtering mega chemical spaces into experimental-ready compound sets using static docking, flexible docking, ensemble docking, Boltz2, and diversity-driven selection.
# End-to-End Virtual Screening Campaign Guide
**Filtering Mega Chemical Spaces into Experimental-Ready Compound Sets**
This guide provides a complete, engine-by-engine protocol for running a computational drug discovery virtual screening campaign — from initial library ingestion through final compound selection for wet lab testing. It covers static docking, flexible docking, ensemble docking, Boltz2 co-folding, chemical space analysis, and diversity-driven final selection.
Follow the funnel in order. Each phase feeds the next. Do not skip phases.
***
## 1. Campaign Philosophy & Funnel Architecture
Virtual screening campaigns operate as a staged triage funnel. The fundamental trade-off at every stage is compute speed vs. accuracy. Faster, lower-accuracy methods process large libraries to remove obvious non-binders, while slower, higher-accuracy methods are reserved for the smaller, pre-enriched pools that survive each cut.
### The Three-Phase Funnel
* **Phase 1 — Static Docking:** Process millions of compounds cheaply. Hard-filter on binding affinity. Carry forward the top 2–5%.
* **Phase 2 — Flexible Docking:** Introduce receptor flexibility + CNN rescoring. Multi-dimensional filter with physical plausibility gates. Output \~3,000–4,000 compounds.
* **Phase 3 — Ensemble Docking:** Use molecular dynamics-derived protein conformers. Aggregate across snapshots. CNN affinity + pose filters. Output \~500 compounds.
**Alternative Funnel** (if benchmarks are operating well; use when in the regime of 3–10k compounds that still require filtration):
* **Phase 4 (optional)** — Boltz2 Co-folding: Deepest structural predictions. Gate on TM score, iPTM, pIC50 or affinity probability. Output top 50–100 candidates.
### Master Decision Summary — Three-Phase Docking Pipeline
| Engine | Library Size | Primary Metric | Key Threshold | Output / Next Stage |
| ---------------- | -------------------- | ------------------------- | ---------------------------- | --------------------------- |
| Static Docking | Full library (1M–2M) | Best Affinity (kcal/mol) | \< −8 kcal/mol | Top 2–5% → Flexible Docking |
| Flexible Docking | \~40–50k filtered | CNN Affinity (desc.) | Intramol ≤ 0; CNN Pose ≥ 0.6 | Top 3–4k → Ensemble Docking |
| Ensemble Docking | 3–5k | CNN Affinity (aggregated) | 80th/20th percentile filters | Top 500 → MD/FEP |
### Alternative Pipeline
| Engine | Library Size | Primary Metric | Key Threshold | Output / Next Stage |
| ----------------- | ------------ | ---------------------- | -------------------- | -------------------------------- |
| Boltz2 Co-folding | 500–1,000 | pIC50 / affinity\_prob | TM ≥ 0.5; iPTM > 0.6 | Top 50–100 → Diversity + Wet Lab |
Library sizes above are representative; scale thresholds based on your compute budget and target class. Calibration R² always drives engine selection — see Section 3.
***
## 2. Pre-Screening: Library Preparation
### 2.1 Source & Deduplication
Before any docking begins, prepare the compound library to ensure clean, unique chemistry.
* Consolidate all library sources (sub-batches, vendors, internal plates) into a single file.
* Deduplicate on **canonical SMILES string** — not on compound identifier or name.
* Record: total input count, unique count, duplicate count. This becomes your baseline.
* Log the deduplication statistics explicitly. They matter for tracking compound attrition across the funnel.
Deduplication by SMILES ensures unique chemistry evaluation. Duplicates skew frequency-based rankings and waste compute.
### 2.2 ADMET & Physicochemical Pre-Filtering (Optional Pre-Screen)
For very large libraries (>500k), applying lightweight physicochemical filters before docking can reduce noise and runtime. These are not required but are recommended when the library source is broad (e.g., a general commercial collection rather than a focused set).
* **Lipinski Ro5:** MW ≤ 500, logP ≤ 5, H-bond donors ≤ 5, H-bond acceptors ≤ 10
* **QED score** (Quantitative Estimate of Drug-likeness): QED ≥ 0.4 for lead-like filtering
* **TPSA:** ≤ 140 Ų (general oral bioavailability proxy)
* **Rotatable bonds:** ≤ 10 (conformational complexity filter)
* **Pan-assay interference compounds (PAINS):** flag and optionally remove reactive, promiscuous scaffolds
You can use **RevADMET** for this task.
Apply these only if your library is chemically unfiltered. For targeted collections (e.g., Enamine US/UA stock), these filters add marginal value and may remove valid hits.
### 2.3 Protein Target Preparation
The quality of the protein structure drives the quality of every downstream docking run. Do not skip this step.
* Source a high-resolution crystal structure for your target (from PDB or AlphaFold2/3 if no crystal exists).
* Remove all co-crystallized ligands, water molecules, and ions that are not part of the binding site.
* Retain metal cofactors that are known to participate in binding (e.g., zinc in coordination sites).
* Verify the intended chain(s) are present; strip extraneous chains not needed for the screen.
* Define the docking box: center on the known binding site or active site. A **30 × 30 × 30 Å box** is typical for most pockets; adjust based on pocket size.
* Run a short MD simulation (10 ns minimum) to check protein stability in solution before running ensemble docking downstream. Monitor RMSD, RMSF, and radius of gyration for convergence.
***
## 3. Calibration Strategy
Calibration is the most important step that most campaigns skip. Before screening thousands or millions of compounds, you must establish which scoring function actually correlates with experimental potency for your specific target. R² from IC50 calibration dictates every engine and readout selection decision downstream.
### 3.1 What to Calibrate
* Use a set of compounds with known experimental activity (EC50, IC50, Ki) against your target.
* 10–20 diverse compounds with at least a 10-fold range in potency is the minimum. Wider is better.
* Include both active and inactive compounds if available.
* Run this calibration set through every docking modality you plan to use (static, flexible, ensemble, Boltz2).
### 3.2 Calibration Metrics to Compare
| Method / Readout | Interpretation |
| --------------------------------- | ------------------------------------------- |
| Ensemble Docking — CNN Affinity | Highest accuracy; preferred primary readout |
| Ensemble Docking — CNN Pose Score | Strong geometry validation |
| Flexible Docking — CNN Affinity | Good fallback; used when ensemble not run |
| Ensemble Docking — Best Affinity | Moderate; supplement with CNN metrics |
| Flexible Docking — Best Affinity | Weaker; use only as supporting signal |
| Boltz2 — pIC50 / log₁₀(IC50) | Exploratory; gate with TM score / iPTM |
| Boltz2 — Affinity Probability | Weak; use as tiebreaker only |
Utilize Pearson Correlations, Spearman Coefficients, and RMSE/MAE to calibrate. Calibrations should be done based on the number of compounds available on your test sets as well (n sensitivity).
* **R² > 0.7** is a reliable readout for ranking — use as primary score.
* **R² 0.4–0.7** is a usable supporting signal — combine with a stronger readout.
* **R² \< 0.4** should not be used as a standalone primary metric.
Always re-run calibration on your specific target — values will vary by protein class, binding site character, and library chemistry.
### 3.3 Calibration Decision Rule
* **Aim for R² > 0.8:** Calibrate across static, flexible, ensemble docking, and co-folding. If you get to >0.8 Pearson, use that as your primary guide for downstream filtering.
* **If R² > 0.6:** Attempt to utilize other engines that represent the biology better, like ensemble docking.
* **If primary calibration metrics are weak:** Fall back to affinity\_probability as tiebreaker, gated by TM score and iPTM.
Calibration is not a one-time activity. Re-calibrate whenever you change the binding site definition, protein model, or add flexibility to new residues.
***
## 4. Phase 1 — Static (Rigid) Docking
Static docking treats both the protein and ligand as rigid bodies. It is the fastest method, making it the only practical choice for screening libraries in the millions. The purpose here is rapid triage: eliminate obvious non-binders, not to find the best poses. It is driven by GPUs so it is able to move much faster.
### 4.1 Setup Parameters
* **Protein:** rigid receptor, prepared as described in Section 2.3
* **Ligand:** rigid SMILES input; do not generate flexible conformers at this stage
* **Exhaustiveness:** 8 for full library triage; increase to 16–32 for smaller batches if time allows; 200 only for calibration compounds
* **Docking box:** 30 × 30 × 30 Å centered on binding site (adjust per target)
* **Batch processing:** split large libraries into batches of 100k–250k for parallelization and fault tolerance
### 4.2 Key Metrics & Thresholds
| Metric | Threshold | Notes |
| ------------------------------ | --------------------------------- | ---------------------------------------------------------------------------------------------- |
| Binding Affinity | \< −8 kcal/mol (strict) | Hard cutoff. Compounds weaker than −8 kcal/mol are deprioritized for downstream runs |
| Binding Affinity — strong hits | \< −10 to −15 kcal/mol | Compounds in this range should be forwarded preferentially to flexible/ensemble docking |
| Number of poses | ≥ 9 poses generated | Inspect top 3 poses for pharmacophoric match to known binding residues |
| Exhaustiveness | 8 (screening) → 200 (calibration) | Use low exhaustiveness for full library triage; increase to 200 only for calibration compounds |
### 4.3 Filtering Logic
1. Remove all rows with missing or null affinity scores.
2. Sort by Best Affinity ascending (most negative = strongest binding).
3. Apply hard cutoff: **Best Affinity \< −8 kcal/mol**.
4. From the passing compounds, take the **top 2–5% by affinity** for the next phase.
5. For targets where known active compounds cluster at −10 to −15 kcal/mol, set the threshold accordingly.
**Distribution sanity check:**
* Compute mean, median, and standard deviation of Best Affinity across the full screened set.
* If mean affinity is weaker than −8 kcal/mol, your library may not contain quality binders for this pocket, or the pocket definition needs adjustment.
* Validate that your known calibration hits fall in the top 5–10% of the distribution.
Static docking will produce false positives. The purpose of this stage is speed-based enrichment only. All static docking hits must be validated through flexible or ensemble docking.
### 4.4 Chemical Space Check (Post Phase 1)
After extracting your top compounds, plot a UMAP or PCA of the filtered set using Morgan fingerprints (ECFP4, radius 2, 2048 bits). Verify:
* The shortlisted compounds are not all clustered in one scaffold region (confirms chemical diversity in your carry-forward set).
* Known active compounds, if available, fall within or near the dense regions of the top-scoring set.
* Isolated outliers in chemical space are not artifacts of the library (check their raw docking scores).
* You can sample across the distribution of chemical spaces to take a more diverse set into further screening. At the end of the day, you are triaging all of these into the wet lab to get primary SAR before lead optimization, so more chemical diversity helps to get more shots on target with diverse chemistries before optimizing within constrained spaces.
***
## 5. Phase 2 — Flexible Docking with CNN Rescoring
Flexible docking allows specified protein side chains (residues in the binding site) to move during the docking calculation, and incorporates a Convolutional Neural Network (CNN) to re-score poses based on geometric and energetic realism. This dramatically improves accuracy over static docking at moderate compute cost.
### 5.1 Setup Parameters
* **Protein:** same structure as Phase 1, but with designated flexible residues enabled
* **Flexible residues:** select residues with known pharmacophoric roles (from co-crystal data, mutagenesis, or MD RMSF analysis). Typically 2–5 residues. Do not make the entire protein flexible.
* **CNN re-scoring:** enabled; adds a geometry-aware neural network on top of classical docking scoring
* **Exhaustiveness:** 8 (default for flexible screen); increase to 32 for highest-priority batches
* **Input:** top-scoring compounds from Phase 1 static screen
### 5.2 Outputs Produced
* **Best Affinity (kcal/mol):** classical empirical binding energy; lower = more favorable
* **Best Intramol (kcal/mol):** intramolecular strain energy of the ligand in the docked pose; higher = more strained
* **Best CNN Pose Score (0–1):** CNN-based assessment of pose geometry and physical realism; higher = more realistic
* **Best CNN Affinity:** CNN-predicted binding affinity (pK metric); higher = stronger predicted binding
### 5.3 Multi-Dimensional Filtering Pipeline
The flexible docking filter is a sequential QC funnel, not a single cutoff. Apply in order:
| Filter / Gate | Threshold | Purpose |
| --------------------- | ---------------------------- | ---------------------------------------------------------- |
| Intramol Strain Veto | Best Intramol ≤ 0 kcal/mol | Removes physically strained (implausible) conformations |
| CNN Pose Quality Gate | Best CNN Pose Score ≥ 0.6 | Validates geometric realism of the docked pose (0–1 scale) |
| Binding Energy Floor | Best Affinity ≤ −10 kcal/mol | Ensures minimum thermodynamic favorability for the pocket |
| Primary Ranking | Best CNN Affinity (desc.) | Final sort: highest CNN affinity compounds advance |
**Final output selection:**
* After all gates pass, sort by Best CNN Affinity descending.
* Select top N compounds (typically 3,000–4,000 as input to ensemble docking).
* Retain a backup pool (top 4,000) if downstream ensemble docking yields insufficient hits.
CNN Pose Score ≥ 0.6 is a geometry threshold, not a strict binary. If your target class shows systematically lower pose scores (e.g., allosteric or shallow sites), adjust downward — but document the change.
### 5.4 Metric Definitions Reference
* **Best Affinity:** Additive, empirical energy calculation. Reflects thermodynamic favorability of the interaction. Lower (more negative) is better.
* **Best Intramol:** Minimum intramolecular energy of the ligand in its docked conformation. Values > 0 indicate steric clashes or physically impossible geometries. Acts as a hard veto.
* **Best CNN Pose Score:** Binary-style geometric validation (0–1 scale) assessing whether the pose looks physically realistic based on thousands of known crystal structures. Values ≥ 0.6 indicate plausible poses.
* **Best CNN Affinity:** Neural network-predicted binding affinity derived from pose geometry and energetics. This is the highest-quality readout from flexible docking and the primary ranking signal.
***
## 6. Phase 3 — Ensemble Docking
Ensemble docking accounts for protein conformational dynamics by docking compounds against multiple representative protein structures, each from a different point in a molecular dynamics trajectory. This captures the protein's natural flexibility beyond individual side chains and substantially reduces false positives.
### 6.1 Generating the Protein Ensemble
* Run an MD simulation of the apo or holo protein for at least 10 ns (100 ns preferred for full equilibration).
* Confirm equilibration: RMSD should plateau; RMSF should show stable core with defined flexible loops; radius of gyration should level off.
* Extract representative snapshots at regular intervals (e.g., every 2 ns for a 10 ns simulation = 5 conformers; every 10 ns for 100 ns = 10 conformers).
* Optionally exclude outlier snapshots where known binding site geometry is disrupted.
* Run docking against each conformer independently, then aggregate scores per compound.
### 6.2 Score Aggregation per Compound
For each unique compound (identified by SMILES), aggregate across all conformers:
* **Best CNN Affinity** → take the maximum across all conformers
* **Best Affinity** → take the maximum across all conformers
* **Best Intramol** → take the minimum across all conformers
This aggregation captures the best observed interaction of the compound with any accessible protein conformation. It is more informative than any single-conformer score.
### 6.3 Filtering Pipeline
| Step | Operation | Rationale |
| ----------------------- | ------------------------------------------------------------------- | -------------------------------------------------------------------------------------- |
| 1. Dedup / Aggregate | Per SMILES: CNN Affinity → max; Best Affinity → max; Intramol → min | Collapses multiple poses per compound to single representative scores |
| 2. CNN Affinity Filter | ≥ 80th percentile of set | Selects top-scoring compounds by neural network affinity prediction |
| 3. Best Affinity Filter | ≤ 20th percentile (within subset) | Confirms thermodynamic favorability via classical scoring within the CNN-filtered pool |
| 4. Intramol Veto | Best Intramol \< 0 kcal/mol | Hard exclusion of strained structures |
| 5. Final Rank | CNN Affinity (desc.) | Sort and take top N |
Percentile-based thresholds (80th/20th) are relative to the screened set. Re-compute percentiles after aggregation, not from pre-aggregation raw scores.
### 6.4 Why Ensemble Docking Outperforms Flexible Docking
* Flexible docking moves only designated side chains. Ensemble docking samples backbone movements and global conformational states that flexible docking cannot reach.
* CNN Affinity R² typically improves by 15–25% going from flexible to ensemble docking in a well-calibrated system.
* Ensemble docking is significantly more compute-intensive. Reserve it for the pre-filtered pool (3k–5k compounds) from Phase 2, not the full library.
***
## 7. Alternative Pipeline — Boltz2 Co-folding
Boltz2 is a structure prediction model that co-folds a protein–ligand complex from sequence and SMILES, generating predicted binding poses and associated confidence metrics. It is the most computationally expensive per-compound method and is reserved for the top 500–1,000 candidates from ensemble docking.
### 7.1 Key Outputs from Boltz2
| Readout | Good Range / Threshold | Interpretation / Notes |
| -------------------- | ----------------------------- | -------------------------------------------------------------------------------------------------------------- |
| pIC50 | > 5.0 (i.e. IC50 \< 10 μM) | Predicted potency in log scale. Use as primary ranking when IC50 calibration R² > 0.5 |
| Affinity Probability | > 0.5 (higher = better) | Model confidence in a binding event. Use as tiebreaker or when IC50 calibration is weak |
| Predicted TM Score | ≥ 0.5 | Structural reliability of the co-folded complex. Gate: reject compounds below threshold regardless of affinity |
| iPTM (interface pTM) | > 0.6 preferred | Interface quality score. High iPTM with low TM = good binding pose but uncertain overall fold; still useful |
| Confidence Score | Use only as supporting signal | Low predictive correlation with experimental IC50 on its own; context-dependent |
### 7.2 Ranking Strategy — Choosing the Right Readout
Use the calibration R² from Section 3 to select your primary readout:
* If IC50 calibration **R² ≥ 0.5 for pIC50:** rank by pIC50 descending; gate by TM score ≥ 0.5 and iPTM > 0.6
* If IC50 calibration **R² \< 0.5:** use affinity\_probability as tiebreaker, with TM score and iPTM as mandatory gates
* In all cases: compute `predicted_ln(ic50)_nM = log₁₀(predicted_ic50_nM)` and include in export for downstream reference
* Always apply confidence gates before using any potency ranking — a high pIC50 with a low TM score is not a trustworthy prediction
### 7.3 Gated Confidence Filtering
* **Gate 1 (Hard):** TM Score ≥ 0.5. Predictions below this threshold indicate the model failed to produce a reliable fold and should be excluded.
* **Gate 2 (Soft):** iPTM > 0.6. Interface quality. Values below this suggest the binding interface is not well-modeled; flag but do not necessarily exclude if other metrics are strong.
* **Gate 3 (Context-dependent):** Affinity Probability > 0.5 as a supporting filter when potency data is unavailable.
Boltz2 calibrates better for some target classes than others. Co-folding performance is weakest for large allosteric sites, covalent binders, and metal-coordinated ligands. Use with appropriate skepticism and weight against docking data.
### 7.4 Hybrid Boltz2 + Ensemble Approach (Exploratory)
As an exploratory check, you can merge Boltz2 outputs with ensemble docking scores:
* Apply the same ensemble pose filters (Section 6.3) to the merged set.
* Rank by predicted\_pic50 descending.
* Use this as a comparison list, not as the primary deliverable.
* If Boltz2 IC50 calibration is valid, the Boltz2-primary ranking supersedes the hybrid for final compound selection.
***
## 8. Chemical Space Analysis
Chemical space visualization is performed at two key points in the campaign: (1) after Phase 1 to verify diversity of the carry-forward pool, and (2) before final compound selection to ensure the final set covers the activity landscape. Skip this step and you risk selecting 50 structurally identical compounds with one structural scaffold, which doesn't help you get the outputs you need for a primary screen — which is primarily SAR.
### 8.1 Fingerprinting
All chemical space analysis uses Morgan fingerprints (ECFP) as the molecular representation:
* **Morgan ECFP4:** radius 2, 2048-bit vector. Standard default.
* **Morgan ECFP6:** radius 3 for finer resolution of large, complex libraries.
* Generate fingerprints for the full screened set and the shortlisted subset simultaneously for direct comparison.
### 8.2 Dimensionality Reduction Methods
| Method | What It Captures | When to Use | Interpretation Tips |
| ------ | ------------------------------------------------------ | ----------------------------------------------------------- | -------------------------------------------------------------------------------- |
| PCA | Global structural diversity (variance-maximizing axes) | First pass: gauge overall chemical space breadth | Tight clusters = structurally similar compounds; spread = diverse library |
| tSNE | Local neighborhoods and cluster relationships | When you need to identify sub-families or scaffold clusters | Not comparable across runs; don't read into global distances |
| UMAP | Both global and local structure simultaneously | Default for visualizing screened sets; best middle ground | Clusters with high CNN affinity/color encoding = activity hotspots to prioritize |
**Recommended workflow:**
1. Run UMAP as your default visualization. It balances global structure and local clusters.
2. Overlay activity scores (CNN affinity, binding energy) as a color dimension to identify activity hotspots.
3. Run PCA as a secondary check to validate that diversity metrics are not UMAP artifact-driven.
4. Run tSNE only when you need to investigate specific scaffold families or local cluster composition.
### 8.3 What to Look For
**Good diversity indicators:**
* Top-scoring compounds (color-coded) are spread across multiple UMAP regions, not concentrated in one cluster.
* High-affinity regions have some overlap with known active scaffolds but also extend into novel chemical space.
* The shortlisted set (e.g., top 3k from flexible docking) covers most of the dense regions of the full filtered pool.
**Red flags:**
* > 70% of top compounds fall within one tight UMAP cluster — indicates scaffold enrichment, not diverse chemistry.
* Known active scaffolds are completely absent from the top-scoring set — may indicate a pocket definition problem.
* The screened set collapses into a single dense region after filtering — likely because a narrow binding energy range was used; consider relaxing the threshold.
### 8.4 Centralized Clusters Correlated with High Activity
It is normal and expected for some chemical clusters to correspond to high CNN affinity or binding activity. This does not disqualify them. The goal is not to eliminate clusters, but to ensure that your final compound list is not exclusively drawn from one cluster while missing other potentially active regions. A compound list of 50 from one scaffold provides weak SAR; 50 compounds spanning 5–10 distinct scaffolds provides a robust SAR foundation.
***
## 9. Final Diversity-Driven Compound Selection
After all computational filters have been applied, you will typically have 100–500 compounds that pass potency and quality thresholds. Final selection to the experimental batch size (typically 50–100) requires balancing hit quality with chemical diversity.
### 9.1 Seed Set Definition
Start by anchoring the selection to confirmed high-quality compounds:
* Extract the top 10 compounds by highest CNN affinity (or pIC50 if Boltz2 was primary).
* Extract the top 10 compounds by lowest binding energy (most negative Best Affinity).
* Deduplicate to produce a **seed set**. This ensures the final list retains the best-scoring compounds regardless of diversity algorithm outcome.
### 9.2 MaxMin Selection
MaxMin (Maximum Minimum Distance) is the standard diversity selection algorithm for compound libraries. It maximizes structural spread across the chemical space.
1. Compute Morgan ECFP4 fingerprints for all candidate compounds.
2. Compute pairwise Tanimoto distance = 1 − Tanimoto similarity for all compound pairs.
3. Initialize with the seed set (best-scoring compounds from 9.1).
4. Iteratively add the compound that is farthest from all currently selected compounds (maximizes minimum pairwise distance).
5. Continue until target selection size is reached.
**Why MaxMin works here:** By seeding with the top-affinity compounds, you guarantee the best hits are included. MaxMin then fills the remainder with maximally diverse chemistry, ensuring coverage of activity hotspots across the chemical space rather than oversampling one scaffold family.
### 9.3 Verification of Final Set
Before finalizing, verify the selected set using UMAP and PCA:
* Plot seed compounds (blue), MaxMin-selected compounds (red), and the full candidate pool (grey).
* Confirm seed and MaxMin compounds are distributed across the major UMAP clusters.
* Verify no single cluster accounts for more than 30–40% of the final selected set (unless target biology demands scaffold specificity).
* If diversity is inadequate, relax the seed set size or broaden the input candidate pool.
Chemical diversity in the experimental set is not just aesthetically preferable — it directly determines the SAR value of the data returned from the wet lab. Redundant scaffolds give you redundant data.
***
## 10. Post-Selection: MD Simulation Validation
The top 50–100 compounds from the final selection should undergo individual molecular dynamics simulation before ordering. This step provides the highest-confidence computational validation and filters the list once more before incurring synthesis or procurement costs.
### 10.1 Short MD Run (0.1 ns) — Screening Mode
* Run all final candidates at 0.1 ns in complex with the target.
* Compute MMPBSA/MMGBSA binding free energy as an approximation.
* Note: 0.1 ns simulations are pre-equilibrium. Energies will be overestimates but are useful for relative ranking and early flag-raising.
* Flag any compound showing immediate ligand displacement or catastrophic pose degradation.
### 10.2 Full MD Run (1–10 ns) — Final Validation
* Take the top 20–30 compounds from the 0.1 ns screen and run for 1–10 ns.
* **Monitor RMSD** of the ligand in the binding site: stable RMSD \< 2–3 Å indicates retained binding.
* **Monitor RMSF** of key binding residues: unexpected rigidification or large fluctuations signal poor complementarity.
* **SASA** (Solvent Accessible Surface Area): ligand binding should reduce solvent exposure in the pocket.
* Compute final MMPBSA/MMGBSA energies. Reference: compounds with −15 to −20 kcal/mol binding free energies are strong candidates.
### 10.3 Binding Free Energy Benchmarks
| Potency Category | IC50 Range | MMPBSA Range |
| ------------------------ | --------------- | ------------------- |
| Tight binders | Sub-100 nM IC50 | −15 to −20 kcal/mol |
| Moderate binders | 1–10 μM IC50 | −8 to −15 kcal/mol |
| Weak / threshold binders | > 10 μM | \< −8 kcal/mol |
Proceed with weak/threshold binders only with strong computational evidence from other engines.
***
## 11. Complete Campaign Checklist
### Pre-Screening
### Calibration
### Phase 1 — Static Docking
### Phase 2 — Flexible Docking
### Phase 3 — Ensemble Docking
### Phase 4 — Boltz2 (Optional)
### Final Selection
Your final experimental compound set should represent diverse scaffolds, pass physical plausibility filters, show strong predicted binding energies across multiple engines, and include both the best-scoring hits and structurally distinct representatives from each major activity cluster.
***
## 12. Glossary
| Term | Definition |
| ------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Best Affinity (kcal/mol) | Classical empirical binding energy from docking. Lower (more negative) = stronger predicted binding. |
| Best Intramol (kcal/mol) | Intramolecular strain energy of the ligand in its docked conformation. Values > 0 indicate physically implausible geometry (hard veto). |
| CNN Affinity | Convolutional neural network-predicted binding affinity. Trained on crystal structures; captures geometric and energetic features beyond classical scoring. |
| CNN Pose Score (0–1) | Binary-style geometric validation of the docked pose. ≥ 0.6 = physically plausible geometry. |
| Exhaustiveness | Docking algorithm search parameter. Higher = more thorough search but slower. Use 8 for triage; 200 for calibration. |
| pIC50 | Negative log₁₀ of IC50 in molar units. Higher = more potent. pIC50 = 9 corresponds to IC50 = 1 nM. |
| TM Score | Template Modeling score from Boltz2. Measures structural similarity of predicted complex to a plausible reference. ≥ 0.5 indicates reliable fold prediction. |
| iPTM | Interface predicted TM score. Measures quality of the binding interface specifically. > 0.6 preferred. |
| Affinity Probability | Boltz2 model confidence in a binding event occurring. Supplements pIC50 as a tiebreaker. |
| MaxMin Selection | Diversity selection algorithm. Iteratively adds compounds maximally distant from the current set by Tanimoto distance. |
| Tanimoto Similarity | Fingerprint-based pairwise molecular similarity metric (0–1). Tanimoto Distance = 1 − similarity. |
| Morgan / ECFP | Extended Connectivity FingerPrint. Circular fingerprint encoding chemical neighborhoods around each atom. ECFP4 (radius 2) is standard. |
| UMAP | Uniform Manifold Approximation and Projection. Dimensionality reduction for chemical space visualization capturing both global and local structure. |
| MMPBSA/MMGBSA | Molecular Mechanics Poisson–Boltzmann / Generalized Born Surface Area. Free energy calculation methods for estimating binding affinity from MD trajectories. |
| R² (Calibration) | Coefficient of determination between predicted docking scores and experimental IC50 values. Primary metric for engine selection. |
| SMILES | Simplified Molecular Input Line Entry System. String representation of molecular structure. Used as canonical identifier for deduplication. |
# High Throughput Screening
Source: https://docs.revilico.bio/high-throughput-screening
A complete walkthrough of structure-based drug design and computational HTS on the Revilico platform — from target identification to a confirmed lead set.
## Overview
A computational high-throughput screening (HTS) campaign on the Revilico platform follows five sequential phases. Each phase feeds the next, so work through them in order on your first campaign.
| Phase | Name | What Happens |
| ----- | ----------------------------------- | --------------------------------------------------------------------- |
| 1 | Target Identification | Identify your biological target and define the therapeutic hypothesis |
| 2 | Structure Acquisition & Preparation | Obtain and clean a high-quality protein structure |
| 3 | The Four Docking Engines | Understand the tools available and when to use each |
| 4 | Calibration | Benchmark each engine against known experimental data |
| 5 | Production Run | Screen your compound library and refine to a confirmed lead set |
***
## Phase 1 — Target Identification
Before running a single computation, you need a clearly defined biological target and a therapeutic hypothesis. This phase is about narrowing the scientific and commercial landscape down to one specific protein — and knowing exactly how you intend to modulate it.
### How Targets Are Selected
Target selection typically draws from three sources:
* **Literature** — peer-reviewed publications, structural genomics datasets, or recently solved crystal structures that reveal a druggable binding site.
* **Multi-omics data** — genomics, transcriptomics, proteomics, or metabolomics evidence that a given protein is causally linked to the disease phenotype of interest.
* **Market intelligence** — financial records, competitive pipelines, and projected indication markets that indicate an unmet therapeutic need and a viable commercial opportunity.
### Define Your Therapeutic Hypothesis
Once you have a target, decide explicitly how you want to modulate it before moving forward.
* **Mechanism** — Are you aiming for competitive inhibition, allosteric modulation, activation, or protein-protein interaction (PPI) disruption?
* **Binding site** — Identify the pocket you want to engage. For a kinase this is typically the ATP-binding site (hinge residue and gatekeeper). For a PPI, identify the hot-spot residues that anchor the interface.
* **Inhibitor type** — Type I (DFG-in), Type II (DFG-out), Type III (allosteric), or covalent. This determines which conformational state of the protein you dock into.
A poorly defined therapeutic hypothesis leads to poor pose interpretation and misguided hit selection. Invest time here before running any computational screen.
***
## Phase 2 — Structure Acquisition & Preparation
Your docking engines are only as good as the structure you feed them. A contaminated or incomplete structure will produce misleading results regardless of which engine you run.
### Step 1 — Obtain the Structure
Source your structure from one of three routes:
Search by protein name, UniProt accession, or indication. Prefer high-resolution crystal structures (below 2.5 Å resolution). A co-crystal structure with a known inhibitor defines the active binding conformation and gives you key contact residues.
Confirm the canonical sequence and link out to deposited structures. Pay attention to isoforms and active/inactive state annotations.
If no experimental structure exists, generate one from the amino acid sequence using AlphaFold, OpenFold, or Boltz-2 co-folding. This gives a clean structure with no co-crystal contamination to remove.
### Step 2 — Clean the Structure
PDB structures almost always contain components that must be removed before docking. Open the file in **PyMOL** or in **RevBench** and remove all of the following:
* **Co-crystal ligands** — any bound inhibitor, agonist, substrate, or cofactor already present in the structure.
* **Water molecules** — crystallographic waters should be removed for standard docking campaigns.
* **Ions and buffer components** — sodium, chloride, DMSO, glycerol, and all other non-protein atoms.
* **Excess chains** — if the structure is a dimer, trimer, or higher-order complex, keep only the chain you want to dock into (usually Chain A). Remove all others.
**Example:** You download a kinase structure containing a staurosporine co-crystal ligand, 180 water molecules, two sulfate ions, and four chains (A, B, C, D). Keep only Chain A and delete everything else. The resulting file should contain only the protein backbone and side-chain atoms of your target domain.
### Step 3 — Retain Co-Crystal Contact Information
Before deleting a co-crystal ligand, record the key residues it contacts. This contact map is invaluable for downstream pocket definition and flexible-residue selection — and lets you validate your later docking poses against known biology.
### Step 4 — Define the Binding Pocket
Use **RevPocket**, or extract the pocket from the literature or from your co-crystal contact analysis, to define a docking box around the site of interest. Verify that the box encompasses the residues identified as mechanistically critical in Phase 1. The box position and dimensions are passed directly to the docking engine.
***
## Phase 3 — The Four Docking Engines
Revilico provides four complementary engines. They are used sequentially in a production run — faster and broader at the start, slower and more accurate at the end.
| Engine | Throughput | Key Outputs | Best Used For |
| -------------------- | -------------- | ----------------------------------------------------------------- | ------------------------------------------------------------------- |
| Co-Folding (Boltz-2) | Per target | Predicted complex structure | Novel targets with no crystal structure; protein-ligand co-folding |
| Static Docking | 2M+ compounds | Binding energy (kcal/mol) | High-throughput initial screen; fast GPU-accelerated elimination |
| Flexible Docking | Up to 50K | Binding energy, CNN affinity, CNN accuracy, intramolecular energy | Hit refinement; binding-site flexibility and steric clash detection |
| Ensemble Docking | \~3K compounds | Binding affinity across MD snapshots | High-confidence lead confirmation; captures protein dynamics |
### Co-Folding (Boltz-2)
Co-folding takes your SMILES string and your protein amino acid sequence as two inputs and generates a predicted complex structure as output. Use this engine when no experimental structure exists, or when you want to evaluate a chemotype in a flexible, co-folded conformation rather than a rigid crystal structure. Most commonly used for novel targets or targets where the binding site is poorly defined.
### Static Docking
Static docking is a GPU-accelerated, semi-empirical scoring method. The protein structure is held completely rigid and the ligand is sampled across thousands of orientations within the docking box. Because the protein does not flex, this engine is very fast — it can handle libraries of 2 million compounds or more.
**Primary output:** binding energy in kcal/mol. More negative values indicate stronger predicted binding.
Use static docking for the initial wide-net screen. Its role is to eliminate clearly non-binding compounds quickly, not to generate publication-quality poses.
### Flexible Docking
Flexible docking extends static docking by allowing up to 3–5 selected protein residues to move during scoring. This is more computationally expensive, so throughput drops to approximately 50,000 compounds or below.
Outputs from flexible docking:
* **Binding energy (kcal/mol)** — same semi-empirical scoring as static docking.
* **CNN affinity** — a pKi value estimated by a convolutional neural network trained on protein-ligand data. Provides an AI-based heuristic for binding strength that complements the semi-empirical score.
* **CNN accuracy** — a confidence score on the predicted pose. Use this to flag unreliable poses quickly.
* **Intramolecular energy** — measures internal strain within the ligand pose. Highly positive values indicate steric clashes within the compound itself, which typically invalidate the pose.
Select flexible residues that you know engage the ligand based on your co-crystal contact analysis or the literature. Poor residue selection undermines the accuracy advantage of this engine.
### Ensemble Docking
Ensemble docking is the most rigorous and computationally intensive method. Before running it, you must first complete a protein-in-water molecular dynamics (MD) simulation using **RevMD Aqua**. Set the simulation time to 100 ns (minimum 50 ns acceptable). The MD run captures how the protein naturally moves in solution.
Pull snapshots at regular time intervals from the completed MD run.
Align the protein snapshots to a common reference frame.
Define the binding site on each snapshot.
Each compound is scored across the full ensemble of protein conformations simultaneously.
Ensemble docking provides binding affinity estimates that account for protein flexibility and dynamics. It is the most reliable predictor of binding in cases where the protein is conformationally dynamic. Use it for your final \~3,000 compound set as the last confirmation step before nominating compounds for wet-lab synthesis.
***
## Phase 4 — Calibration
Calibration is the single most important quality-control step in any computational screening campaign. Without it, you have no basis for trusting the scores your engines produce. Calibration tells you how well each engine predicts experimental activity for compounds with known data — and therefore how much confidence to place in its predictions for unknowns.
### Build a Calibration Set
Collect a set of compounds with known experimental activity against your target. Aim for several hundred to several thousand compounds with associated IC₅₀ or Kᵢ values. Biochemical assay data is preferred over cell-based.
Highly curated with good metadata. Generally trustworthy as a primary source.
Broad coverage. Verify assay annotations carefully before including.
Acceptable, but audit assay conditions and metadata. Inconsistent protocols between sources introduce noise.
Use **RevBench** to scrape and compile these compounds. Export a CSV with at minimum two columns: SMILES and experimental activity value (with units and assay type annotated). The compounds should be chemically diverse and representative of the chemical space you intend to screen.
### Run All Four Engines on the Calibration Set
Take the calibration library and run it through each of the four engines exactly as you would run a production screen. Use the same pocket box and settings you defined in Phase 2. Collect all outputs into a master CSV: one row per compound, one column per engine output.
### Generate Calibration Plots and Interpret Metrics
In RevBench, upload your combined CSV alongside the experimental activity data and generate calibration plots — predicted score vs. experimental value — for each engine. Evaluate the following four metrics:
| Metric | Range | Interpretation |
| -------------------- | ---------------------- | -------------------------------------------------------------------------------------------------------------------------------------------- |
| Pearson / R² | 0–1 (higher is better) | Measures linear correlation between predicted and experimental values. R² > 0.5 is a reasonable baseline for virtual screening. |
| Spearman Coefficient | 0–1 (higher is better) | Measures rank-order correlation. More important than R² in screening, where ranking compounds correctly matters more than absolute accuracy. |
| RMSE | Lower is better | Root Mean Squared Error. Sensitive to outliers. Use alongside MAE to understand whether a few bad predictions are skewing the average. |
| MAE | Lower is better | Mean Absolute Error. A robust average error metric less sensitive to extreme outliers than RMSE. |
No engine will be perfect across all assay types. Look for which engine performs best for your specific target class. Compare speed vs. accuracy trade-offs. The calibration data gives you a defensible, data-driven rationale for your engine choices in the production run.
***
## Phase 5 — Production Run
With calibration complete and engine performance understood, you are ready to screen a real compound library. The production run is a staged funnel — each stage reduces the compound count while increasing scoring accuracy.
| Stage | Engine | Compound Count | Goal |
| --------------------- | ---------------- | -------------- | --------------------------------------------------- |
| 1 — Library Screen | Static Docking | \~2,000,000 | Eliminate non-binders rapidly |
| 2 — Hit Refinement | Flexible Docking | 30,000–50,000 | Score and rank filtered hits with improved accuracy |
| 3 — Lead Confirmation | Ensemble Docking | \~3,000 | Validate binding across protein conformations |
### Stage 1 — Static Docking Screen (2M+ Compounds)
Select your compound library. Common sources include the Enamine REAL Library, Enamine 2M liquid stock compounds, or any other commercially available or in-house library. Upload the SMILES to RevBench as a CSV.
For libraries of 2 million or more compounds, **split the library into 10–15 equal batches**. Submit all batches simultaneously — they run in parallel and are automatically concatenated once complete.
After the run completes, apply a binding energy cutoff:
* Binding energy more negative than a defined kcal/mol threshold (e.g., more negative than the best-performing calibration compounds, or a fixed threshold such as −8 kcal/mol).
Then apply secondary filters:
* **Chemical space clustering** — remove redundant chemotypes to ensure diversity in your hit set.
* **Binding affinity ranking** — sort by score and apply a top-N cutoff.
**Target output:** 30,000–50,000 compounds carried forward to Stage 2.
### Stage 2 — Flexible Docking (30K–50K Compounds)
Import your filtered hit set into the flexible docking engine. Select 3–5 flexible residues at the binding site based on your calibration analysis and co-crystal contact data. Run the batch — again split into 10–15 parallel jobs if needed.
Evaluate each pose using all four flexible docking outputs:
* **Binding energy** for a primary rank.
* **CNN affinity** to cross-check with the AI-based prediction.
* **CNN accuracy** to flag low-confidence poses for manual review.
* **Intramolecular energy** to remove poses with internal steric clashes.
**Target output:** \~3,000 high-confidence hits carried forward to Stage 3.
### Stage 3 — Ensemble Docking (\~3,000 Compounds)
Run the protein-in-water MD simulation in RevMD Aqua at 100 ns if you have not done so already. Extract trajectory snapshots at regular intervals. Load the snapshots into the ensemble docking pipeline, define the pocket on each snapshot, and submit your 3,000-compound set. Split into batches of 10–15 for parallelism.
Ensemble docking produces per-compound binding affinity scores averaged across all protein conformations captured in the simulation. Compounds that score well consistently across many snapshots are the most likely to bind in a real, dynamic biological environment.
Apply a final round of filtering and ranking. The resulting top compound set is your lead series — ready for molecular dynamics validation and prioritization for wet-lab synthesis.
After Stage 3, you will have a ranked lead set with high-confidence binding poses, affinity predictions from multiple orthogonal methods, and validated alignment to a biologically relevant binding site. This set is now ready for physical synthesis and biochemical assay confirmation.
***
## Quick Reference Checklist
Use this checklist to track your campaign progress.
**Phase 1 — Target Identification**
* Identify target from literature, multi-omics, or market analysis
* Define the mechanism: inhibition type, binding site, key residues
**Phase 2 — Structure Preparation**
* Download structure from PDB / UniProt, or generate with AlphaFold / Boltz-2
* Remove all co-crystal ligands, waters, ions, and excess chains
* Record co-crystal contact residues before deletion
* Define and validate the binding pocket using RevPocket or literature
**Phase 3 — Engine Selection**
* Review calibration data to understand engine performance for your target class
* Confirm 100 ns MD simulation is queued for ensemble docking stage
**Phase 4 — Calibration**
* Compile calibration set (ChEMBL / PubChem / literature, biochemical assay preferred)
* Run all four engines on calibration set with production settings
* Generate calibration plots; record R², Spearman, RMSE, MAE per engine
* Select best engine for Stage 1 screen based on speed vs. accuracy
**Phase 5 — Production Run**
* Stage 1: Static docking on 2M+ compounds, split into 10–15 batches
* Apply binding energy cutoff and chemical diversity filter; reduce to 30K–50K
* Stage 2: Flexible docking on 30K–50K; evaluate all four output metrics
* Stage 3: Ensemble docking on \~3K; use MD snapshots from RevMD Aqua
* Rank final lead set; nominate for MD validation and wet-lab synthesis
# Revilico Platform
Source: https://docs.revilico.bio/index
Revilico is digitizing pharmaceutical R&D, empowering every chemist and biologist with unified virtual simulations and closed loop experimentation to develop better medicines faster, and with greater precision.
The End-to-End Drug Discovery Platform
## Learning Resources
Step-by-step workflow guides covering high-throughput screening, platform
best practices and advanced use cases
Hands-on video walkthroughs for docking, co-folding, structure prediction,
target analysis and more
Latest platform updates, new engine releases, performance improvements and
bug fixes
## Platform Engines
Multi-omics and transcriptomics for target discovery
Structural biology, pocket detection and PTM analysis
Docking, MD simulations and generative chemistry
Quantum chemistry and compound property prediction
Virtual cell assays and biological perturbation models
Experimental orchestration (Coming Soon)
Autonomous AI agent-driven research workflows
3D visualization, compound design and data management
## RevOmics
Multi-omics target identification using single-cell transcriptomics,
epigenomics, and genetic risk analysis.
scRNA-seq clustering, UMAP and cell type annotation
DNA methylation and epigenomics profiling
Polygenic risk scoring, GWAS and genetic variant analysis
## RevTarget
Structural analysis, druggable binding site detection, and post-translational
modification profiling.
Predict 3D protein structure from amino acid sequence
Open-source structure prediction and protein folding
Druggable binding site detection and allosteric pocket search
Phosphorylation site prediction and kinase substrate mapping
## RevBind
The full computational binding toolkit: molecular docking, MD simulations,
cheminformatics analytics, and generative chemistry.
RevDock: Molecular Docking and Co-folding
Static, flexible and ensemble virtual screening at scale
Protein-ligand complex structure prediction via co-folding
Generative structure-based binder design
RevAnalytics: Cheminformatics and SAR
ML-powered QSAR and activity prediction models
3D pharmacophore modeling and shape-based screening
Docking protocol benchmarking and scoring evaluation
RevSynth: Generative Chemistry
De novo molecular library generation with generative AI
Multi-objective lead optimization with reinforcement learning
Custom generative model fine-tuning for focused libraries
RevDynamics: Molecular Dynamics Simulations
Protein-water MD for stability and conformational sampling
Protein-ligand MD, binding kinetics and MM-GBSA
Ligand-membrane permeability MD for GPCR targets
RevFEP: Free Energy Perturbation
Relative and absolute binding free energy (RBFE / ABFE)
Protein mutation stability and binding affinity change (PMX)
## RevQuant
Quantum chemistry, ADMET profiling, retrosynthetic planning, and quantum
simulation across eight specialized tools.
DFT geometry optimization and energy minimization
HOMO-LUMO and frontier orbital analysis
Aqueous solubility and solvation prediction
Full ADMET profiling: PK, toxicity and drug-likeness
Low-energy 3D conformational ensemble generation
AI retrosynthetic route planning and synthetic accessibility
Quantum simulation for binding energy and protein folding
Transition state search and reaction activation energy
## RevSim
Virtual cell biology: in silico perturbation models and high-content assay
simulation across the full drug discovery pipeline.
RevPerturb: Cell Perturbation Modeling
Virtual IC50 and dose-response cell viability prediction
Gene regulatory network modeling and perturbation analysis
Perturbation sequencing: single-cell CRISPR screen analysis and causal GRN
inference
RevAssay: Virtual Assay Suite
Apoptosis, necrosis and caspase cascade prediction
IL-6, TNF-a and cytokine storm immunotoxicity profiling
mTOR, MAPK and PI3K downstream signaling analysis
Cancer cell motility, wound healing and transwell invasion
Growth rate inhibition metrics and GR50 antiproliferative analysis
## RevLab: Experimental Orchestration
These experimental orchestration modules are in active development and will
launch soon.
Enzyme inhibition, fluorescence and luminescence-based biochemical screening
automation
High-content cell-based phenotypic imaging and in vitro screening workflows
Automated synthesis routing, compound ordering and ELN integration
## RevIntel
An autonomous AI agent that orchestrates multi-step drug discovery workflows
across all Revilico engines, synthesizing findings, generating hypotheses, and
producing structured research reports from target identification through to
lead candidates.
## RevStudio
Design, visualize, and manage your research data across the entire drug
discovery pipeline.
Interactive 3D molecular and structural visualization
Structure-based compound design and editing workspace
Multi-property compound triage and review interface
AI-powered literature and patent scanning
RevData: Data Management
Centralized cloud storage for all experimental and computational data
Collaborative data editing, annotation and review workflows
Structured scientific notebooks for reproducible research
## Solutions
End-to-end workflows combining multiple Revilico engines to solve real-world
drug discovery challenges, from first hypothesis to IND-ready candidate.
Multi-omics analysis, structural hypotheses, pocket druggability and
therapeutic cell line selection
Virtual screening, binding mechanism resolution and compound library
prioritization
Scaffold expansion, multi-parameter optimization and metabolite-inspired
design
Deep characterization, ADMET profiling, toxicity risk and synthetic route
planning
Cross-team collaboration, reproducibility, version control and pipeline
transparency
# Therapeutic Strategy and Hit Matching
Source: https://docs.revilico.bio/solutions/hit-identification/Identify-Hits-That-Match-a-Specific-Binding-Modality
Identify Hits That Match a Specific Binding Modality or Mechanism
## **The Problem You Are Trying to Solve**
*“I have a biological target, and I want to identify a set of hits that meet a specific mechanistic or chemical criterion, like specific modes of action: inhibition, activation, covalent binding, or other specialized mechanisms.”*
At hit identification, not all hits are equal. Different discovery goals require fundamentally different binding behaviors, and a generic virtual screen often fails to distinguish them. Common challenges include:
* Docking scores that don’t reflect functional outcome (inhibition vs activation)
* Difficulty identifying covalent or mechanism-based binders
* Binding poses that look plausible but don’t explain functional modulation
* Mixing multiple binding modalities in a single hit list
This workflow is designed to explicitly tailor hit identification to your intended mechanism of action, producing hits that are aligned with *how* you want the target to be modulated.
## **Solution**
This workflow uses Revilico’s binding chemistry engines to customize hit discovery based on binding modality, combining structural filtering, interaction analysis, and (when needed) dynamic or energetic validation. The shared core workflow is: Target Preparation → Mechanism-Aware Docking → Interaction & Motif Analysis → Hit Prioritization.
Where the workflow diverges is in:
* How docking is configured
* What interactions are emphasized
* Which downstream validation steps are applied
Integration with MD, FEP, QSAR, and generative chemistry is possible once modality-specific hits are identified, and the campaign has reached an iterative optimization stage.
## **What Data Do I Need to Provide?**
Required
* Target protein structure (experimental or predicted)
* Compound library (or libraries) to screen
Recommended
* Known ligands or reference compounds for the desired modality
* Information about active sites, allosteric sites, or reactive residues
Optional
* Cofactors, ions, or partner proteins
* Experimental benchmarks for validation
## **Workflow**
1. **Define the Desired Binding Modality**
Before screening, explicitly define *what kind of hit you want*.\
Common criteria include:
* Inhibitors (competitive, non-competitive)
* Activators / positive allosteric modulators
* Covalent binders
* Allosteric binders
* Stabilizers / disruptors of protein–protein interactions
* Membrane-active or permeability-driven binders (when relevant)
This decision informs every downstream configuration choice to ensure that binding and engagement are translating to biologically relevant functions.
2. **Prepare Target Structures Based on Modality**
Different modalities require different structural contexts.
* Active-site binders (inhibitors):
* Prepare catalytic or orthosteric site
* Use known substrate or inhibitor coordinates when available as a reference
* Allosteric modulators / activators:
* Identify non-canonical pockets via literature, pocket detection, or prior MD
* Multiple conformations may be needed
* Covalent binders:
* Identify reactive residues (e.g., Cys, Ser, Lys)
* Ensure correct protonation and residue accessibility
Primary engines used
* **Docking** (for pose hypothesis)
* Optional **Protein Water MD** (to reveal transient or cryptic pockets) with potential follow up by the Pocket Search Engine to identify pockets of interest.
This step will give you one or more docking-ready target representations aligned with the mechanism.
3. **Mechanism-Aware Virtual Screening**
#### **A. Inhibitor Discovery**
Primary engines
* **Static Docking** → **Flexible Docking** (for refinement)
Key focus
* Competitive occupation of known binding pocket
* Strong, well-oriented interactions with catalytic residues
* Consistent low-energy poses
What you prioritize
* Affinity trends, especially hits that are reading out lower/negative kcal/mol energies
* Pose stability, looking at physically feasible binding and lack of steric clashes
* Overlap with known inhibitor binding modes, attaching to key amino acids of interest in the pocket
#### **B. Activators / Allosteric Modulators**
Primary engines
* **Docking** (alternative pockets)
* **Pharmacophore Analysis**
* **Molecular Dynamics Simulation Protein Ligand and Protein Water for Calibrations**
Key focus
* Binding outside the active site
* Interactions that stabilize specific protein conformations
* Poses that plausibly alter protein dynamics, and can be measured through molecular dynamics and Root Mean Squared Fluctuation (RMSF) values
What you prioritize
* Pocket specificity, ensuring that you are engaging with the right allosteric pocket
* Interaction networks rather than raw affinity scores, and physically feasible conformations
* Consistency across protein conformations
Validating the results
* In order to validate that your designed compound is an allosteric modulator, you can put the pose into molecular dynamics simulation over longer time span and look for RMSF value increases across longer time scales, in the active site of the given protein.
* Your designed compound should reveal a larger fluctuation pattern in the protein after ligand engagement compared to just protein in water MD
#### **C. Covalent Binder Identification**
Primary engines
* **Docking** (pose feasibility)
* **Pharmacophore Analysis** (reactive geometry validation)
Key focus
* Proper orientation of electrophile toward nucleophilic residue
* Feasible reaction geometry (distance, angle)
* Avoidance of nonspecific reactivity as covalent binders have a tendency to have high non-specific binding to other molecules in the body, causing a higher potential of toxicity.
What you prioritize
* Geometry over docking score
* Accessibility and selectivity of the reactive residue
*Integration note:* Protein–Ligand MD can later validate pre-reactive stability before covalent bond formation.
#### **D. Protein–Protein Interaction (PPI) Modulators**
Primary engines
* **Docking** (large, shallow interfaces)
* **BoltzGen Co-Folding**
* **Pharmacophore Analysis**
**Key focus**
* Hotspot engagement, targeting key residues at the PPI interface that contribute the most to binding free energy
* Surface complementarity ensuring that you have geometric and electrostatic fit between the ligands and the protein interface surface
* Disruption or stabilization of interface contacts where the compound could either disrupt the native protein protein contacts, or stabilize them like a molecular glue between proteins.
**What you prioritize**
* Shape complementarity along the interface
* Multi-point interaction patterns, ensuring engagement across the interfaces
* Interface-specific binding rather than engagements in other regions of the protein
4. **Interaction Pattern Analysis**
After screening, abstract away from individual poses. Use **Pharmacophore Analysis** to:
* Identify interaction motifs shared by top hits
* Confirm alignment with desired mechanism
* Filter out compounds that bind “correctly” but not *usefully*
This step ensures hits are mechanistically meaningful, not just high-scoring. The Pharmacophore Engine also will enable you to iterate on key motifs on the molecule with functional groups to optimize binding to key residues in the pocket.
5. **Dynamic or Energetic Validation (Optional)**
For high-confidence or high-cost decisions, escalate selectively. Depending on modality:
* **Protein–Ligand MD** to validate pose stability and conformational effects along with a breakdown of the energies using snapshot Free Energy calculations
* **Ligand–Membrane MD** for permeability- or membrane-driven mechanisms for protein targets on membrane interfaces
* **ABFE / RBFE** to refine ranking among closely related candidates within narrower chemical spaces.
This step is not required for all campaigns but strengthens confidence where needed.
## **Results**
* A modality-specific hit list aligned with your biological objective
* Structural explanations for *how* compounds engage the target
* Reduced false positives from generic screening
* Clear mechanistic rationale for experimental follow-up
## **Now What?** *I have hits matched to my desired mechanism, but what’s next?*
Typical next steps include:
* Focused experimental validation
* Mechanism-specific hit expansion
* Multi-parameter optimization
* Transition into lead optimization workflows
## **Why Revilico?**
Revilico enables mechanism-aware hit identification by allowing users to:
* Explicitly encode intent (inhibitor vs activator vs covalent, etc.)
* Configure screening and analysis engines accordingly
* Preserve structural interpretability throughout the workflow
* Seamlessly escalate into deeper validation or optimization pipelines
This ensures that hits are not just binders, but binders that behave the way you need them to.
# Screening Compound Libraries
Source: https://docs.revilico.bio/solutions/hit-identification/Prioritize-a-Compound-Library
Prioritize a Compound Library Against a Target Using Virtual Screening
## **The Problem You Are Trying to Solve**
*“I have a biological target and a large compound library that would be expensive and time-consuming to screen experimentally. I want to computationally prioritize a smaller, higher-confidence subset of compounds before committing to HTS.”*
At the hit identification stage, experimental HTS can be:
* Costly and slow at large library scales
* Noisy, with high false-positive and false-negative rates
* Difficult to iterate on quickly
This workflow is designed to front-load computational triage, allowing you to focus experimental resources on compounds most likely to bind and be biologically relevant.
## **Solution**
This workflow uses Revilico’s Virtual Screening and Binding Chemistry engines to down-select large libraries into a ranked, mechanistically interpretable shortlist.
The primary prioritization chain is: Target Preparation → Virtual Screening (Docking) → Pose & Score Analysis → Optional Refinement (Flexible / Ensemble Docking).
This approach allows users to:
* Rapidly screen large libraries
* Eliminate obvious non-binders
* Preserve structural insight into *why* compounds were prioritized using structure activity relationships and chemical space analysis.
* Seamlessly escalate promising candidates into deeper simulations if needed
Other engines (QSAR, MD, FEP, generative chemistry) can be layered on later as confidence or optimization needs increase.
## **What Data Do I Need to Provide?**
Required
* Target protein structure (experimental or predicted)
* Compound library (CSV with SMILES strings), or if you’d like access to our partner’s 2M liquid stock libraries for direct delivery after computational screening, reach out to us.
Recommended
* Known binding site or reference ligand (to define docking region)
* Any known cofactors, ions, or biologically relevant states of the target
Optional
* Multiple protein conformations (for flexible or ensemble docking)
* Experimental benchmark compounds (for calibration and validation).This usually will consist of ligand sets with corresponding experimentally determined activity values.
## **Workflow**
1. **Prepare the Target for Screening**
Before screening, ensure the target structure is suitable for docking. On Revilico, users typically:
* Upload or retrieve a protein structure (PDB or predicted model)
* Clean the structure (remove waters, ions, irrelevant ligands)
* Define the binding site (known ligand coordinates, residue-based centroid, or pocket detection)
If target flexibility or uncertainty is expected, this step can later integrate **Protein Water MD** to generate alternate conformations, and to assess for certain vulnerabilities static algorithms may possess for certain protein target systems..
This results in a docking-ready protein structure with a defined binding region. If you are still assessing the proper protein pocket, you can utilize our pocket search engine after generating structures using AlphaFold, Boltz2, or OpenFold.
2. **Rapidly Screen the Full Library with Static Docking**
Begin with high-throughput prioritization. This step prioritizes speed and coverage, not perfect accuracy.
Use **Static Docking** to:
* Screen tens of thousands to millions of compounds efficiently using GPU scaled docking algorithms
* Predict binding poses and approximate affinities
* Quickly eliminate compounds with poor shape or interaction complementarity
What you’re looking for
* Strong predicted affinities relative to the bulk library
* Plausible poses that occupy the intended binding pocket
* Consistency across multiple poses
This gives a ranked list of compounds with predicted binding scores and poses. After the initial screening utilizing static docking, the library will still have uncertainty in the predicted binding affinities as these algorithm types tend to have higher rates of false positives and negatives, necessitating a deeper screen with the following algorithms listed below.
3. **Triage and Filter Docking Results**
Refine the initial hit list using structural and statistical filters. This step ensures that prioritization is not driven by docking noise alone.
Users typically:
* Apply affinity cutoffs. The more negative the activity, the better the target engagement.
* Remove compounds with unstable or highly strained poses
* Inspect pose clustering to avoid single-pose artifacts
* Preserve chemical diversity while downselecting
This filtering step outputs a reduced, higher-quality candidate set suitable for deeper analysis.
4. **Refine Binding Predictions with Flexible Docking**
For the most promising compounds, increase physical realism by utilizing the flexible docking engine which allows the protein side chain to be flexible on certain selected residues.
Use **Flexible Docking** to:
* Allow selected protein side chains to move
* Capture induced-fit effects missed by rigid docking
* Re-rank compounds based on improved pose accuracy
* Re-score poses generated with convolutional neural network (CNN) filters, to get better pose accuracies
This step is especially valuable when:
* The binding site is flexible, and critical amino acids for binding are known
* Small chemical differences need better resolution to differentiate activity cliffs within narrower chemical spaces
This results in a refined ranking with higher confidence in binding modes.
5. **Account for Protein Dynamics with Ensemble Docking (Optional, but Highly Recommended)**
If the target is highly dynamic or known to adopt multiple binding-competent states:
* Use **Protein Water MD** to generate conformational snapshots, and to get refined parameters quantifying the protein’s behaviors in different solvents and time scales.
* Apply **Ensemble Docking** across these structures to assess target engagement over the course of the protein’s trajectory in solution.
This captures:
* Pocket breathing
* Transient sub-pockets
* Conformational selection effects of ligand engagement
This will result in prioritized compounds that bind consistently across protein states.
## **Results**
* A ranked, prioritized compound subset suitable for experimental testing
* Structural explanations for prioritization decisions
* Reduced HTS cost and time by focusing on high-value candidates
* Clear upgrade path into hit validation and optimization workflows
## **Now What?** *I have a prioritized list, but what’s next?*
Common next steps include:
* Experimental HTS or focused biochemical assays
* Binding mechanism analysis (MD, pharmacophore analysis)
* Hit expansion or optimization using generative chemistry
* Energetic validation with MMPBSA or FEP for top candidates
## **Why Revilico?**
Revilico enables cost-effective hit identification by combining:
* High-throughput **Virtual Screening**
* Physically grounded refinement (**Flexible / Ensemble Docking**)
* Transparent structural and energetic interpretation
* Seamless escalation into deeper simulation or optimization workflows
This allows teams to screen smarter, not bigger, dramatically reducing experimental burden while increasing the likelihood of meaningful hits. Like all other engines, properly calibrating your screens with some sort of experimental data subset will increase the interpretability of the results and allow for more accurate filtering criterias.
# Elucidating Binding Mechanisms
Source: https://docs.revilico.bio/solutions/hit-identification/Resolve-Binding-Mechanisms-and-Molecular-Interactions
Resolve Binding Mechanisms and Molecular Interactions from Experimental Activity Data
## **The Problem You Are Trying to Solve**
*“I have experimental binding or activity data, and I want to understand why these molecules behave the way they do: what interactions are driving binding, what mechanisms explain activity differences, and how this informs next design decisions.”*
At the hit identification stage, experimental data confirms *something is happening*, but the mechanism is often unclear. Common challenges include:
* Multiple plausible binding modes for the same ligand
* Activity trends that are difficult to rationalize from chemistry alone
* False positives or ambiguous hits without structural explanation
* Limited intuition on which interactions are essential versus incidental
This workflow is designed to translate experimental measurements into mechanistic, structure-level understanding that can guide confident next steps.
## **Solution**
This workflow uses Revilico’s binding chemistry engines to map experimental observations onto molecular interactions, combining static structure analysis, dynamic simulations, and energetic decomposition.
The primary analysis chain is: **Virtual Screening/Docking → Pharmacophore Analysis → Protein–Ligand MD →** Energetic Decomposition (**MMPBSA / ABFE** where needed)**.**
This creates a layered interpretation:
* Docking proposes *how* molecules could bind, and with what magnitude
* Pharmacophore analysis identifies *what interactions matter,* and provide a foundation to engineer new scaffold variants in lead optimization
* MD validates *whether those interactions persist dynamically* over longer time scales and in different conditions
* Energetic analysis explains *why binding is favorable or unfavorable*
Other engines (QSAR, co-folding, generative chemistry) can integrate downstream once mechanisms are clarified. These other engines allow for clustering of the chemical space for hot spots of activity, more rigorous pose and affinity analysis, and for the generation of novel scaffolds to be tested later on.
## **What Data Do I Need to Provide?**
Required
* Experimental binding or activity data (IC₅₀, Kᵢ, Kᴅ, etc.)
* Chemical structures of tested molecules (SMILES or SDF)
* Target protein structure (experimental or predicted)
Recommended
* A subset of representative compounds across activity ranges (strong, moderate, weak)
* Any known binding site information or reference ligands
Optional
* Mutagenesis data or SAR trends
* Known cofactors, ions, or binding partners
* Membrane context (if target or ligand behavior suggests it matters)
## **Workflow**
1. **Anchor Experimental Data to Structural Hypotheses**
Start by proposing plausible binding modes that could explain your experimental results.\
Use the **Virtual Screening Engine** (Static,Flexible, or Ensemble) to:
* Generate candidate binding poses for active and inactive compounds
* Compare pose consistency across potency ranges
* Identify conserved versus variable interactions
At this stage, you are not optimizing scores, you are asking:
* *Do more active compounds share a common binding geometry?*
* *Do weak binders fail to make key contacts or show unstable poses?*
This step gives a set of plausible binding hypotheses grounded in experimental activity.
2. **Identify Key Interaction Motifs**
Next, abstract away from individual poses to identify interaction-level patterns. Use **Pharmacophore Analysis** to:
* Extract hydrogen bond donors/acceptors, hydrophobic features, aromatic interactions, and charge centers
* Compare pharmacophores across active vs inactive compounds
* Identify interaction features that correlate with experimental activity
This step answers:
* *Which interactions appear necessary for activity?*
* *Which regions tolerate variation?*
* *Are there missing interactions explaining weak activity?*
* *What potential modifications can be made for later stage lead series expansion?*
This step results in a mechanistic interaction hypothesis explaining experimental trends.
3. **Validate Binding Mechanisms Under Dynamics**
Static poses alone cannot capture binding stability or induced fit effects.\
Use **Protein-Ligand MD** to:
* Test whether docked poses remain stable over time
* Observe interaction persistence (H-bonds, salt bridges, hydrophobic contacts)
* Identify conformational rearrangements or ligand drift
* Compare dynamics between high- and low-activity compounds
This step helps to distinguish true binders from docking artifacts and allows for confirmation of stable binding modes from transient contacts generated.
This engine outputs a dynamic validation of binding mechanisms with residue-level insight.
4. **Decompose Energetic Drivers of Binding**
Once stable binding modes are established, quantify *why* binding is favorable, and what are the primary thermodynamic drivers causing strong/weak interactions.\
Use:
* **MMPBSA / MMGBSA** for fast energetic breakdown across MD trajectories
* **ABFE** selectively when absolute binding favorability must be quantified
These engines help to:
* Decompose van der Waals, electrostatic, and solvation contributions, among other energetic components.
* Identify residues or interactions dominating binding energetics, and how these change over time scales
* Explain experimental rank ordering in energetic terms
This results in an energetic rationale that complements experimental measurements, to create a better picture of the system being analyzed.
5. **Cross-Validate with Data-Driven Trends (Optional)**
Generalize insights across a broader dataset.
Use **QSAR Modeling** to:
* Identify structure–activity correlations and hot spots of activity within chemical space
* Validate whether interaction hypotheses scale across chemical space
* Flag outliers or inconsistent data points
This step is especially useful when:
* Experimental datasets are large
* Multiple binding modes may exist
## **Results**
* A clear, mechanistic explanation of experimental binding/activity data
* Identified key residues and interaction motifs driving activity
* Dynamic validation of binding hypotheses
* Energetic decomposition supporting observed trends
* A defensible model of *how* and *why* molecules bind
## **Now What?** *I understand the binding mechanism, but what’s next?*
Typical next steps include:
* Designing new compounds that reinforce key interactions
* Eliminating false positives before hit expansion
* Feeding insights into generative chemistry on the platform using Molecular Optimization or Scaffold Decoration engines to create new molecular hypotheses
* Prioritizing compounds for synthesis or further biophysical assays
Now that you have an understanding of the mechanistic drivers of your interactions, you can have a greater chance of engineering the right molecules using the foundational knowledge determined in this workflow.
## **Why Revilico?**
Revilico enables experimental data interpretation by connecting:
* Structural hypotheses (**Docking**)
* Interaction abstraction (**Pharmacophore Analysis**)
* Physical realism (**Protein–Ligand MD**)
* Quantitative energetics (**MMPBSA / ABFE**)
This layered approach transforms experimental measurements into actionable molecular insight, allowing teams to move from “we have hits” to “we understand why they work.”
# Expand Chemical Libraries
Source: https://docs.revilico.bio/solutions/hit-to-lead/Generate-an-Expanded-Small-Molecule-Library
Generate an Expanded Small-Molecule Library From Hit Compounds
**The Problem You are Trying to Solve:**\
*“I have an identified target and a set of initial hit compounds that bind well, and I want a new library for further testing with an expanded set of chemical hypotheses.”*
This can be difficult because expanding from hits to a broader library requires generating compounds that are new enough to teach you something, but close enough to keep activity, and still reasonable to synthesize and test. In traditional hit-to-lead, this happens through repeated medicinal chemistry cycles (analog enumeration, scaffold hopping, fragment linking), supported by screening and structure-based design. With AI generative chemistry, we can propose thousands to millions of hypotheses quickly, but they must be generated in a controlled and interpretable way to be useful. With AI-driven generative chemistry, we can now rapidly generate, prune, and prioritize large hypothesis libraries, provided outputs are scored, filtered, diversified, and interpreted with the right confidence and constraints. Further testing and scoring can be conducted with a variety of Revilico’s Engines to be able to ensure that designed structures will perform well experimentally before undergoing lengthy synthesis cycles.
**Solution**\
This workflow enables users to generate an expanded compound library using Revilico’s Generative Chemistry Suite, powered by reinforcement learning and a variety of scoring functions. Revilico supports three complementary ways to expand beyond your hit set:
* **De Novo Library Generation**, which allows you to explore new chemical space (broad expansion, scaffold hopping, fragment linking, scaffold decoration)
* **Molecular Optimization** generates improved “next-iteration” analogs of your hits under clear goals (property + activity objectives) for multi-parameter optimization of leads.
* **Custom Model Training** creates a fine-tuned generative model on your chemistry so the libraries match your project’s chemical style and constraints, allowing for optimization within pre-determined chemical spaces.
Most teams use a *mix* of these approaches so they get conservative analogs, moderate exploration, and a small number of high-novelty ideas.
**What Data Do I Need to Provide?**\
Required
* Hit compounds as SMILES (CSV or text; CSV must contain one column labeled ‘smileString’
* Your determination of your “expanded library” (close analogs vs scaffold hops vs both)
* How many new hypotheses you want (hundreds, thousands, millions)
Optional (recommended for stronger control)
* Known scaffolds, fragments, or warheads (if you want constrained generation)
* Preferences or constraints (size, drug-likeness, avoid substructures, etc.)
* Project-specific compound dataset (if you want custom model training to remain optimized towards your own chemical spaces)
**Workflow**
1. **Choose Your Expansion Strategy**
Before you generate, decide what kind of expansion you want from your hits:
* Close-in expansion: “Give me analogs close to my hit series”
* Balanced expansion: “Keep the core ideas, but explore substitutions and related scaffolds”
* Exploratory expansion: “Find new chemotypes / scaffold hops”
* Constrained expansion: “Keep my scaffold or warheads fixed and only vary linkers / R-groups”
This choice determines which engine(s) you run first.
2. **Generate a Broad Expansion Library**
Use this step when you want to quickly explore chemical space and create a first-pass expanded library for downstream screening on different Engines. On Revilico, you select a **De Novo Library Generation** mode based on your intent:
* De Novo Generation (no starting molecules needed): broad ideas from general drug-like space biased towards certain property optimizations
* Scaffold Decoration (you provide a scaffold with \* attachment points): explore R-group combinations while preserving the core
* Fragment Linker Design (you provide two warheads separated by |, with \* attachment points): explore linker chemistry between fragments
Revilico produces a large set of candidate molecules and automatically performs quality checks and diversity handling (e.g., removing duplicates and pruning overly similar structures).
**How to interpret the outputs**\
Each generated molecule is accompanied by an NLL score, which helps you understand how “normal” vs “novel” the molecule is relative to the model’s learned chemistry:
* Lower NLL → safer / more typical chemistry (often good for early hit expansion)
* Higher NLL → more novel chemistry (often useful for scaffold hopping or getting unstuck from difficult performance regimes)
3. **Generate “Better Versions” of Your Hits**
Use this step when you want the next library to be hit-like, but improved, based on the priorities you care about. **Molecular Optimization** takes your hit SMILES and produces new “siblings” optimized toward goals you define, such as:
* staying within drug-like property ranges
* increasing desirability under predicted ADMET or physicochemical constraints
* encouraging novelty without breaking the series
* Optimizing for activity against specific known targets using binding affinity scoring functions
* penalizing known liabilities (reactive groups, toxic motifs)
Revilico runs this as an iterative optimization campaign: generating molecules, scoring them, and refining the generator toward better solutions in the next rounds of optimization.
**What you get**
* a ranked library of optimized hypotheses
* clear scoring summaries so you can understand why a molecule was preferred
* optional diversity controls so the library doesn’t collapse into near-duplicates
4. **Train a Custom Generator on Your Chemistry (Optional)**
Use this step when you want your generated libraries to reflect your organization’s chemical space, and not generic public chemistry. If you provide a curated SMILES dataset, Revilico’s models can train or fine-tune a model with **Custom Model Training** so future libraries:
* match your chemistry patterns
* better respect your synthesis constraints
* produce compounds that “feel native” to the project based on already known structure activity relationships your chemists determined
Once trained, that model becomes a selectable option inside De Novo Generation or Molecular Optimization, so you can generate libraries that are consistently aligned with your internal chemical space, helping you to train your own models for downstream lead optimization campaigns and iterative chemical series expansion.
5. **Consolidate Your Libraries Into a Single “Next Test Set”**
Most users will run two or three generation passes, then combine them:
* Library A (Conservative): close analogs of hit series
* Library B (Balanced): moderate exploration around scaffolds / substitutions
* Library C (Exploratory): a smaller set of scaffold hops or fragment-link ideas
From there, users typically apply a consistent triage strategy (e.g., dedupe, cluster, filter, then send to docking/MD/ADMET) and select a final testable subset.
**Results**
* Versioned expanded compound library (new chemical hypotheses)
* A mixture of conservative and exploratory ideas (depending on your chosen strategy)
* Interpretable outputs that help you reason about novelty vs feasibility
* Libraries ready for downstream screening and prioritization
* When multiple strategies produce overlapping conclusions (e.g., similar motifs across de novo + optimization), you can move forward with greater confidence that you’re expanding in a meaningful direction.
**Now what?** I have a new library, but what’s next?\
This workflow commonly feeds into hit-to-lead prioritization steps such as:
* docking and structure-based triage
* molecular dynamics on top candidates
* ADMET filtering and multi-parameter ranking
* synthesis planning and experiment selection
* Before sending these results to the wet lab, any other Revilico engine can be used to curate compound sets with more optimal properties before spending money on synthesis or testing.
**Why Revilico?**\
Revilico allows you to expand from hits to a next-generation test library using three complementary approaches (broad exploration, guided optimization, and project-specific model training) while keeping outputs organized, versioned, and interpretable. This enables rapid iteration without losing scientific control over what is being generated and why.
# Design Allosteric Modulators
Source: https://docs.revilico.bio/solutions/hit-to-lead/Metabolite-Inspired-Drug-Development
Metabolite-Inspired Drug Development & Allosteric Optimization
## **The Problem You Are Trying to Solve**
*“I have a binder to my protein, but I want to design synergistic or allosteric modulators to improve target engagement and reduce reliance on traditional, trial-and-error SAR campaigns.”*
For some cases, a primary binder (orthosteric ligand), especially with activities in the 1-100uM range, is not enough to:
* Achieve desired potency in complex biological systems
* Improve selectivity over homologous proteins
* Modulate protein function dynamically (activation vs inhibition)
* Enhance stability or residence time
Allosteric modulators and metabolite-inspired designs offer powerful solutions, but experimentally screening for allosteric sites and combinations is time-consuming and expensive. Moreover, elucidating the detailed breakdown of the molecular mechanisms driving these dynamic interactions can be difficult to experimentally resolve.
## **Solution**
This workflow identifies potential allosteric sites, designs modulators inspired by other small molecule engagers like metabolites or secondary binding pockets, and computationally validates synergistic engagement before experimental investment.\
The primary design chain is: Primary Binder Analysis → Allosteric Site Identification → Allosteric Modulator Design → Binding & Stability Validation → Synergy Evaluation.
This enables rational exploration of allosteric control and multi-site engagement with minimal experimental burden.
## **What Data Do I Need to Provide?**
Required
* Protein structure (PDB or predicted structure)
* Primary binder structure (SMILES or PDB complex)
Recommended
* Known compound or metabolite structures (if metabolite-inspired design is desired)
* Known regulatory partners or secondary binders
* Any experimental binding/activity data for calibration
Optional
* Target conformational states of interest (active vs inactive)
* Known mutational or regulatory hotspot regions
## **Workflow**
1. **Analyze the Primary Binder and Binding Mode**
Start by deeply understanding how your existing ligand engages the protein. Users typically:
* Visualize the orthosteric binding pose with **Docking**
* Identify key residues and interaction motifs
* Validate stability and observe dynamic contacts with **Protein-Ligand MD**
* Determine which structural regions remain unoccupied
**Key questions:**
* Which residues drive binding?
* Are there flexible regions or distal pockets that move during MD?
* Does binding induce conformational shifts?
This gives a high-confidence model of orthosteric engagement and dynamic behavior.
2. **Identify Potential Allosteric Sites**
Allosteric sites often emerge through dynamics, flexibility, and conformational coupling. Users typically:
* Run **Protein-Water MD** to observe natural pocket breathing, and feed the simulations into **MDPocket** engine where transient pockets can be elucidated for further analysis.
* Analyze RMSF, SASA, and PCA outputs to identify flexible regulatory regions
* Examine transient pocket formation during MD trajectories
**Indicators of candidate allosteric regions:**
* Transient pockets forming distal to the active site
* Correlated motion between distant domains
* Stable but ligand-free cavities that open during simulation
This produces candidate allosteric pocket(s) for modulator targeting.
3. **Design Allosteric or Metabolite-Inspired Modulators**
Once candidate pockets are identified, design molecules that engage them. Users can:
* Run **Virtual Screening / Docking** on the allosteric pocket
* Use **De novo Library Generation** for new chemical matter
* Use **Molecular Optimization** to decorate metabolite-like scaffolds
* Use **Pharmacophore Analysis** to match pocket features
If metabolite-inspired:
* Upload metabolite SMILES
* Identify shared motifs with endogenous ligands
* Preserve functional groups important for recognition
* An example of this would be kinase inhibitor designs replicating ATP structures that drive engagement for proteins that do downstream phosphorylation
This results in your shortlist of predicted allosteric binders.
4. **Validate Allosteric Binding Stability**
Now confirm that predicted modulators bind stably and plausibly. Users run:
* Initial assessments of the orthosteric ligand in the primary pocket can be done with docking, co-folding, or molecular dynamics simulations
* **ProteinLigand MD** for the allosteric binder alone
* Optionally simulate both orthosteric + allosteric ligand together. **Revilico’s Boltz Co-Folding** Engine allows for you to simulate multiple ligands binding within 1 structure, seeing how activity will change with just 1 ligand binding with the protein versus with both compounds.
* Using downstream Protein Ligand MD for either system will help elucidate pockets opening up on the protein with longer term dynamic time scales.
**What you’re evaluating:**
* Stable binding in the secondary pocket
* No destabilization of protein core
* Conformational shifts induced by allosteric binding
This provides your set of stability-validated allosteric binder candidates.
5. **Evaluate Synergistic Engagement**
The key goal is enhanced target engagement or functional modulation. Users can:
* Simulate dual-bound systems (orthosteric + allosteric ligand) with co-folding with pooled ligand pairs on one specific target and then running the ligands on **Protein–Ligand MD** simulations with long enough time spans to capture protein motion induced with ligands.
* Compare conformational landscapes via **PCA**
* Evaluate binding free energy shifts via:
* **MMPBSA and MMGBSA** (rapid estimate)
* **ABFE/RBFE** (higher rigor with alchemical transformations of the ligand)
**Key evaluation questions:**
* Does the allosteric ligand stabilize the orthosteric binder?
* Does the protein adopt a more favorable active/inactive conformation?
* Is dual engagement thermodynamically favorable?
This quantifies synergy and mechanistic insight for the design of allosteric ligands.
## **Results**
* Identified and validated allosteric pocket(s)
* Designed metabolite-inspired modulators
* Stability-confirmed binding models
* Quantified synergistic target engagement
* Reduced need for brute-force SAR experimentation
## **Integration with Other Engines (Optional)**
This workflow can connect to:
* **QSAR Modeling** (learn patterns from orthosteric + allosteric activity data)
* **ADMET-AI** (ensure new modulators remain developable)
* **Quantum Chemistry (HOMO–LUMO, Geometry Optimization)** for reactivity/stability checks
* **Free Energy Perturbation** for fine analog ranking
* **scRNA-Seq Analysis** to assess downstream pathway effects of modulation
## **Why Revilico?**
Revilico enables metabolite-inspired and allosteric drug design by combining:
* Structural insight (Docking + MD)
* Dynamic conformational analysis
* Generative chemistry capabilities
* Thermodynamic validation (FEP)
* Integrated interpretation tools
Instead of running large, unfocused SAR campaigns, you computationally guide allosteric discovery toward the most mechanistically promising and synergistic designs.
# Multi-Parameter Lead Optimization
Source: https://docs.revilico.bio/solutions/hit-to-lead/Multi-Parameter-Optimize-a-Moderately-Active-Molecule
Multi-Parameter Optimize a Moderately Active Molecule via Scaffold Decoration and Scaffold Hopping
**The Problem You Are Trying to Solve**\
*“I have a moderately active small molecule, and I want to improve it across multiple dimensions (potency, selectivity, developability, etc.) by exploring scaffold decoration and scaffold hopping, while preserving key motifs that drive favorable target interaction.”*
At this stage, you’re no longer in discovery mode, determining what chemistries work for certain use cases, you're deciding how to evolve the chemistry within certain chemical space constraints without breaking what makes the molecules active. This is traditionally difficult because:
* Multi-parameter goals often conflict (e.g., potency vs solubility vs lipophilicity) and optimizing structure for one property usually will do the opposite for other properties you are concerned with
* Scaffold hopping can accidentally remove the very interaction features responsible for binding
* Minor changes can unexpectedly shift binding mode, strain, or developability risk
This workflow is designed to help you explore broader chemistry while staying anchored to the motifs that matter.
**Solution**\
This workflow enables users to run multi-parameter optimization (MPO) using Revilico’s generative chemistry suite, specifically leveraging Scaffold Decoration and Scaffold Transformations / Diversification to propose new analogs and scaffold hops that retain key interaction motifs. The primary engine chain is: QSAR and Pharmacophore Analysis → De Novo Library Generation (Scaffold Decoration / Transformations) → Docking → Molecular Optimization (using a variety ofMPO stages).
This creates a controlled loop where:
* You define what must be preserved (motifs + interaction features)
* You generate scaffold-level hypotheses that respect those constraints
* You triage structurally (docking) and iteratively improve multi-objectives (Molecular Optimization)
You optionally validate top candidates with MD / FEP when confidence needs to be higher
**What Data Do I Need to Provide?**\
Required
* Starting molecule(s) as SMILES (moderately active lead(s))
* Target protein structure (experimental or predicted)
* A clear statement of what “better” means (your MPO priorities)
* You must generally know structure activity relationships of your compounds so you know what scaffolds and motifs you must preserve when running optimizations. This can be done through virtual or experimental screening and co-crystal structure determinations
Recommended
* Known binding pose(s) or docking results for the starting molecule
* A description of the key motifs to preserve (functional group, ring system, H-bond pattern, charge center, etc.)
* Any known constraints or liabilities (substructures to avoid, MW limits, LogP/TPSA ranges, etc.)
Optional
* Experimental activity / selectivity data (improves scoring and QSAR usefulness)
* A set of related analogs (helps define SAR and guide similarity constraints)
**Workflow**
1. **Define the Set to Preserve: Motifs + Interaction Features**
Before exploring new scaffolds, establish what cannot be lost. On Revilico, users typically:
* Run **Docking** (Static or Flexible) on the starting molecule to confirm a plausible binding mode
* Use **Pharmacophore Analysis** to capture the binding-critical interaction features (donor/acceptor patterns, hydrophobics, charge centers, spatial arrangement)
* Identify the motif(s) that should be preserved during decoration/hopping (e.g., hinge binder, charged anchor, aromatic stacking group)
* If experimental data is already obtained at this stage, you can re-run the compounds in docking/pharmacophore engines to gauge target engagement across certain motifs and in QSAR modeling to extract necessary substructures that drive activity.
This step produces a clear “motif + interaction hypothesis” that will guide scaffold decoration/hopping and later scoring decisions.
*Integration note:* If you have historical activity/ADMET data, you can optionally use **QSAR Modeling** here to help identify which structural features correlate with activity or liabilities.
2. **Generate Decorated Analogs Around the Core Scaffold**
Now expand around your existing scaffold while preserving the core. Use **De Novo Library Generation: Scaffold Decoration** to:
* Keep the scaffold fixed
* Specify attachment points
* Generate diverse R-group combinations that explore chemical space around your motif-preserving core
This step is ideal when:
* You believe the core scaffold is correct
* You want to systematically explore substituents to improve MPO dimensions (potency + solubility + stability + etc.)
This step results in a scaffold-preserving analog library that explores R-group space broadly and efficiently.
3. **Perform Scaffold Hopping / Core Replacement While Preserving Motifs**
If your current scaffold is limiting (e.g., poor developability, IP issues, metabolic liabilities), move from decoration into scaffold hopping. Use **De Novo Library Generation: Molecular Optimization** modes (e.g., scaffold-based transformations or broad scaffold diversification) to propose scaffold hops that:
* Retain key interaction motifs (pharmacophore features)
* Explore alternate cores that could improve MPO properties (developability, selectivity, stability)
This step is ideal when:
* You want new chemotypes (not just close analogs)
* You suspect your current core is a liability but the binding motif is correct
This will create a scaffold-hop library: new cores with motif-preserving features and medicinally relevant chemistry.
*Integration note:* If your organization has a proprietary chemical style or modality constraints, this is a great point to optionally use Custom Model Training to bias the generator toward “native” project chemistry.
4. **Re-score for Binding Plausibility**
Now quickly remove structural false positives and prioritize binders, by running **Docking** on your generated library:
* Start with Static Docking if the library is large
* Use Flexible Docking for a smaller set or when binding-site residue movement matters
What you’re looking for
* Plausible binding poses that preserve key interactions
* Strong scores without obvious steric clashes
* Consistency in top poses (not just one lucky pose)
This step will produce a downselected set of motif-consistent, pose-plausible candidates for deeper MPO refinement.
5. **Multi-Parameter Optimization Loop**
With promising scaffold variants in hand, switch from “generate broadly” to “optimize deliberately.” Use **Molecular Optimization Engine** to run MPO by:
* Defining multi-objective scoring (physicochemical bounds + predicted activities + penalties)
* Using staged optimization to sequentially prioritize constraints (e.g., Stage 1: maintain motif + binding plausibility; Stage 2: improve developability; Stage 3: refine potency/selectivity)
This step is where you:
* Balance tradeoffs (potency vs LogP vs PSA vs MW vs liabilities)
* Keep the model anchored to reasonable chemistry (avoid drifting into unrealistic structures)
* Control exploration vs conservatism using similarity settings, diversity penalties, and scoring structure
This engine will give you an MPO-optimized candidate list with transparent scoring breakdowns.
6. **Validate the Best Candidates with Dynamics and Energetics (Optional)**
When you have a small shortlist and need higher confidence before synthesis:
* Use **Protein-Ligand MD** to confirm binding stability and rule out pose artifacts
* Use **ABFE/RBFE Calculation** to get higher-fidelity energetic ranking within close chemical series
This is especially valuable when docking can’t distinguish subtle improvements among similar candidates, and will provide you with a final shortlist supported by structure, stability, and (optionally) binding free energy confidence
**Results**
* A motif-preserving scaffold-decorated and scaffold-hopped candidate set
* MPO-optimized compounds ranked across potency and developability priorities
* A defensible shortlist of improved hypotheses ready for synthesis or experimental validation
* Optional physics-based validation to reduce failure risk downstream
**Now what?** I have MPO-optimized scaffold variants, but what’s next?
* At this point you can continue this cycle until you have a short list of compounds that you’d like to take into synthesis. The chemical hypotheses you have should be rigorously validated computationally.
* For other parameters you’d like to optimize for, you can utilize any other Revilico Engine
**Why Revilico?**\
Revilico makes scaffold-level MPO practical by connecting:
* motif definition (Pharmacophore Analysis)
* broad exploration (Scaffold Decoration + Scaffold Transformations)
* fast structural triage (Docking)
* deliberate multi-objective refinement (Molecular Optimization MPO stages)
* optional high-confidence validation (Protein–Ligand MD, ABFE/RBFE)
This allows teams to explore chemical space aggressively without losing scientific control, and to produce candidates that remain interpretable, testable, and aligned with real drug discovery constraints.
# Optimizing Molecular Properties
Source: https://docs.revilico.bio/solutions/hit-to-lead/Optimize-a-Moderately-Active-Molecule
Optimize a Moderately Active Molecule for Potency, Selectivity, and Developability
**The Problem You Are Trying to Solve**\
“I have a moderately active molecule, and I want to optimize it for higher potency, better selectivity, and improved developability.”
At this stage of hit-to-lead, the goal is no longer just to find binders, but to systematically improve a known chemical series. This is challenging because gains in one dimension (potency) often come at the expense of others (selectivity, solubility, safety, or synthetic feasibility). Traditional optimization relies on slow, iterative medicinal chemistry cycles supported by assays and structural reasoning, which can be costly and time-consuming. With modern AI-assisted workflows, we can propose and evaluate many optimization hypotheses in parallel, but only if generation, scoring, and validation are tightly integrated and interpretable. This allows for multi parameter optimization for every molecule designed and eventually synthesized prior to being sent into the lab.
**Solution**\
This workflow enables users to iteratively optimize a moderately active molecule using Revilico’s integrated generative, analytical, and structure-based engines, with Molecular Optimization as the core driver. At a high level, the workflow:
* Generates improved analogs of the starting molecule
* Scores and prioritizes them using fast, complementary signals
* Validates top candidates with higher-fidelity methods
* Repeats the loop as needed until clear lead candidates emerge
The emphasis is on a controlled optimization loop, rather than one-off generation.
**What Data Do I Need to Provide?**\
Required
* Starting molecule(s) as SMILES (moderately active compounds)
* Target protein structure (experimental or predicted)
* Clear optimization intent (e.g. “increase potency without increasing lipophilicity”)
Optional (highly recommended)
* Known liabilities or constraints (avoid motifs, MW limits, polarity ranges)
* Selectivity context (off-targets or related proteins)
* Property priorities (potency vs PK vs safety tradeoffs)
* Historical activity or ADMET data (for QSAR guidance)
**Workflow**
1. **Establish an Optimization Baseline**
Before generating new molecules, establish what “better” means. On the Revilico Operating System, users typically:
* Review **Static Docking**, **Flexible Docking**, or **Ensemble Docking** results (and/or experimental data) for the starting molecule to understand binding modes and pose stability
* Use **Pharmacophore Analysis** and **QSAR Modeling** to identify key interactions, liabilities, and regions of the molecule that drive activity or risk
* Decide whether optimization should be conservative (close analogs via Molecular Optimization with tight similarity constraints) or more exploratory (relaxed similarity, scaffold or substituent changes)
This baseline informs how aggressively the Molecular Optimization Engine should push the chemistry and which constraints or objectives should guide the optimization loop. Usually, the docking and activity engines will create a baseline for structure activity relationships that can be leveraged during downstream compound generations.
2. **Generate Optimized Analogs**
This is the core step of the workflow. Using Revilico’s **Molecular Optimization** Engine, users input their moderately active molecule(s) and define optimization goals. The engine then generates new “siblings” of the starting compound, iteratively improving them under a multi-objective scoring framework. This can be done through a ‘global iteration’, allowing the engine to modify any portion of the compound, or a more conservative approach of scaffold decoration, preserving core motifs driving activity, while maintaining other side chains or R-groups that can be iterated against. Typical optimization goals include:
* Improving predicted binding or docking performance
* Staying within desirable physicochemical ranges
* Penalizing known liabilities or unstable motifs
* Maintaining similarity to the active series (or relaxing it, if needed)
The engine works by repeatedly:
1. Generating candidate molecules
2. Scoring them against defined objective functions (engines that help predict parameters)
3. Updating the generator model during reinforcement learning to favor better chemistry as predicted by the scoring functions.
This balances exploitation (refining what works) with exploration (avoiding local SAR traps).
This generates a ranked set of optimized molecular hypotheses and clear scoring summaries explaining why molecules were favored.
3. **Rapid Structural and Statistical Triage**
Once a new batch of optimized analogs is generated, users typically apply fast screening layers to prioritize which molecules are worth deeper analysis in other Revilico Engines, depending on what properties need to be optimized (like activity in this case).
**Docking**
* Used to evaluate binding modes and relative affinity trends across static, flexible, and ensemble conditions
* Helps eliminate obvious false positives
* Provides structural intuition for SAR decisions
**QSAR Modeling**
* Uses data-driven patterns to predict activity, selectivity, or developability signals
* Scales well across larger libraries
* Complements docking by capturing non-structural trends
At this stage, the goal is down-selection, not final validation.
4. **Validate Top Candidates with Dynamics and Energetics (Optional)**
For the most promising candidates, users can integrate more expensive but higher-confidence methods.
**Protein-Ligand Molecular Dynamics**
* Tests binding stability over biologically relevant time scales
* Reveals water effects, flexibility, and pose robustness
* Helps eliminate unstable or over-fit docking poses
* This engine is also equipped with snapshot Free Energy Perturbation (FEP) calculations using MMPSA and MMGBSA to get more accurate read outs of energies.
**Free Energy Perturbation (RBFE / ABFE)**
* Provides quantitative ranking within a focused chemical series
* Particularly useful when choosing which compounds to synthesize next
* Often used as a final filter before experimental commitment
* Is capable of highly resolving the breakdown of binding energy contributors for more resolved understanding of ligand protein engagement.
These steps are optional but powerful when optimization decisions become costly, and chemical space becomes more narrow.
5. **Iterate the Optimization Loop**
Based on what you learn:
* Refine scoring objectives
* Adjust similarity constraints
* Introduce new penalties or priorities
* Re-run Molecular Optimization with updated guidance
Most hit-to-lead campaigns cycle through this loop multiple times, progressively narrowing toward a small set of strong lead candidates, overall helping to save a tremendous amount of time and money on synthesized compounds which can eventually fail.
**Results**
* Iteratively improved compound series
* Clear rationale for why each optimization step was taken
* Reduced uncertainty before synthesis or experimental testing
* A small, prioritized set of lead-like candidates ready for the next stage
When improvements are consistently supported across generation, docking, QSAR, and (optionally) physics-based validation, users can proceed with much higher confidence.
**Now what?** I've identified optimized candidates, but what’s next?
* After identifying these candidates and optimizing their activities and other properties using generative chemistry, you can move forward with Retrosynthesis to design and explore synthetic routes to utilize.
* If you’d like to re-analyze the compound set for different key properties that were also flagged for optimization, the rest of the operating system suite can be used for this as well.
**Why Revilico?**\
Revilico enables a closed-loop optimization workflow where molecule generation, scoring, and validation are tightly integrated. Rather than relying on a single signal, users can hedge decisions across generative chemistry, structure-based modeling, and data-driven analytics, allowing optimization to move faster without sacrificing scientific control or interpretability.
# Gaining Confidence in Your Hits
Source: https://docs.revilico.bio/solutions/hit-to-lead/Refine-an-Initial-Hit-Set
Refine an Initial Hit Set into a High-Confidence Subset
**The Problem You Are Trying to Solve**\
*“I have an initial set of hits from a virtual screen, but I want a refined set with fewer false positives and false negatives, validated by more robust, physics-based simulations.”*
Early hit lists are often noisy. Both experimental high throughput screening and virtual screening is designed for speed and scale, which means:
* Some compounds score well for the wrong reasons (false positives)
* Some true binders are missed due to imperfect sampling or rigid assumptions (false negatives)
* Binding stability and solvent effects are not fully captured
This workflow helps you upgrade your given confidence in your hits by progressively applying more realistic models of binding.
**Solution**\
This workflow refines your hit list using a primary refinement chain: Docking → Protein–Ligand MD → ABFE/RBFE. Each step increases realism and confidence for your compounds of interest:
* Docking provides a fast structural sanity check and rescoring method that can be applied to larger library sizes
* Protein-Ligand MD tests whether poses are stable over time in solvent, and with dynamic time conditions as well
* ABFE/RBFE quantifies binding energetics with much higher fidelity than docking
Revilico also supports integration with other engines (QSAR, Pharmacophore Analysis, Protein Water MD, etc.), but the chain above is the core refinement backbone for gaining more confidence in your hit sets, and validating assumptions and hypotheses at the molecular level.
**What Data Do I Need to Provide?**\
Required
* Hit molecules as SMILES (CSV preferred)
* A protein structure (experimental or predicted) with a defined binding site
* A clear binding pocket definition (from known ligand, site annotation, or docking grid)
Optional (but useful)
* Known reference ligand(s) or co-crystal pose (for benchmarking)
* Experimental activity data (if any)
* Notes on liabilities to avoid (reactive groups, solubility risk, etc.)
**Workflow**
1. **Re-score and Sanity-Check Hits**
Start by re-evaluating the hit set with the **Docking** engine to quickly filter obvious false positives. The typical approach is as follows:
* If your hit set is large: run Static Docking first for throughput, then downselect based on set binding criteria for your project
* If your hit set is smaller: go straight to Flexible Docking for better pose realism, or to Ensemble Docking for more advanced workflows and protein systems
What you’re looking for:
* Clear, plausible binding poses (no severe clashes)
* Consistent strong scores across top poses (low variation between best and average)
* A pose that makes chemical sense (key interactions, reasonable geometry)
The Docking engine will output a narrowed hit set with selected docked complexes (protein + ligand poses) ready for MD.
*Integration note:* If the binding pocket is known to be flexible or uncertain, you can optionally generate a receptor ensemble via Protein Water MD and use those snapshots for an “ensemble-style” docking pass before MD.
2. **Validate Binding Pose Stability**
Docking is still a snapshot. Now test whether the hit remains stably bound in a more realistic environment using **Protein-Ligand MD**. In Protein-Ligand MD, your docked complex is solvated, ionized, minimized, equilibrated, then simulated over time, producing a trajectory that shows whether binding is stable or fragile.
What you’re looking for
* RMSD vs time: a quick rise then plateau (no long-term drift)
* Radius of gyration / SASA: stable protein behavior (no unfolding or destabilization)
* RMSF near pocket residues: no extreme instability; stable binding-site behavior
* Visual check: ligand stays in the pocket and does not flip, drift, or dissociate from the pocket
This step will result in a smaller set of hits where binding is stable over time, accompanied by trajectories and stability metrics that help explain “why this hit is real”.
*Integration note:* Protein-Ligand MD can optionally include MMPBSA-style analysis as a faster energetic check, but the main goal here is stability and realism before FEP-level investment.
3. **Quantify Binding Energetics**
Once you have MD-validated hits, move to **ABFE/RBFE Calculation** for a stronger energetic signal. This step helps you answer:
* ABFE: “Is binding thermodynamically favorable, and how strong is it?”
* RBFE: “Between similar ligands, which binds better and by how much?”
Our Free Energy Perturbation Suite uses alchemical free energy methods (lambda windows + MD sampling), accounting for:
* Solvent reorganization
* Protein–ligand dynamics
* Entropy and restraint corrections
* More realistic thermodynamics than docking scores
What you’re looking for
* More favorable (more negative) binding free energies for stronger candidates
* Reasonable, stable energetic components (no obvious artifacts)
* Consistency across methods (TI vs MBAR) when available
* Clear ranking that separates top candidates from the rest
This step will leave you with a high-confidence ranked subset of hits with binding free energy estimates, and a defensible shortlist that is much less sensitive to docking noise.
4. **Final Refinement and Selection**
At this point, you should have a refined hit set that is:
* structurally plausible (Docking)
* dynamically stable (Protein–Ligand MD)
* thermodynamically supported (ABFE/RBFE)
From here, users typically:
* Select a small number of compounds for synthesis / purchase
* Prepare follow-on hit-to-lead expansion campaigns
*Integration note:* You can optionally add **QSAR Modeling** and/or **Pharmacophore Analysis** as additional triage layers (especially if you have assay data), but they are not required for the core physics-based refinement chain.
**Results**
* A refined set of hits with significantly fewer false positives
* Better protection against false negatives (through better pose realism and dynamics)
* Ranked candidates supported by higher-fidelity binding energetics
* Compounds ready for experimental validation or lead optimization workflows
**Now what?**
* You can move these compounds more confidently into experimentation, knowing that these hypotheses have been rigorously tested through a multi-modal procedure to assess for different binding dynamics.
* After being tested initially, you can utilize generative chemistry workflows to optimize your lead sets using similar methods
**Why Revilico?**\
Revilico makes hit refinement a single connected workflow: you can start with virtual screening outputs, validate stability with MD, then quantify binding with FEP-grade calculations, all without rebuilding inputs across tools. This produces a refined shortlist that is both computationally grounded and scientifically interpretable, so teams can make better downstream decisions with less wasted synthesis and assay effort.
# Multi-Plex Tox Screens
Source: https://docs.revilico.bio/solutions/lead-optimization/Assess-Lead-Toxicity-Risk
Assess Lead Toxicity Risk via Target-Based Computational Screening
## **The Problem You Are Trying to Solve**
*“I have a set of lead compounds, and I want to determine whether they are likely to be toxic by screening them against a panel of protein targets known to contribute to toxicity in my indication.”*
At the lead optimization stage, toxicity is one of the most common and costly causes of failure. Experimental tox panels are expensive and slow, and many liabilities only emerge late unless proactively assessed. Typical challenges include:
* Off-target binding that is not obvious from primary efficacy screens
* Toxicity arising from unintended protein interactions (e.g., ion channels, nuclear receptors, metabolic enzymes)
* Difficulty distinguishing *true* toxicity risk from noise in early data
* Lack of mechanistic insight into *why* a compound is toxic
This workflow is designed to computationally triage toxicity risk early, providing both *predictions* and *mechanistic explanations* to guide safer optimization.
## **Solution**
This workflow combines target-based binding screens with AI-driven toxicity prediction to evaluate whether lead compounds are likely to engage known toxicity-associated proteins. The primary toxicity assessment chain is: Toxicity Target Panel Definition → Virtual Screening (Docking) → Interaction Analysis → ADMET-AI Toxicity Prediction → Virtual Cell Line models for Enterprise
This layered approach allows users to:
* Screen leads against known toxicity-relevant proteins
* Identify off-target binding risks before experiments
* Understand *which interactions* may drive toxicity
* Prioritize compounds for safer optimization paths
Other engines (MD, FEP, quantum chemistry, retrosynthesis) can be integrated when deeper validation or redesign is required.
## **What Data Do I Need to Provide?**
Required
* Lead compound structures (SMILES or CSV)
* A curated panel of toxicity-associated protein targets (PDBs or predicted structures)
Recommended
* Known toxicity benchmarks (e.g., hERG binders, CYP inhibitors)
* Experimental safety flags (if available)
Optional
* Multiple conformations or states of toxicity targets
* Membrane context (for ion channels or transporters)
## **Workflow**
1. **Define a Toxicity Target Panel**
Begin by explicitly defining *what toxicity means* for your indication. Typical toxicity panels may include:
* Ion channels (cardiac channels, neuronal channels)
* Metabolic enzymes (CYP450s)
* Nuclear receptors (PXR, CAR, AhR)
* Transporters (P-gp, BCRP)
* Stress-response or apoptosis-related proteins
This step ensures the screen is hypothesis-driven, not generic. This step results in a curated, indication-relevant toxicity target panel.
2. **Prepare Toxicity Targets for Screening**
Each target must be docking-ready and biologically relevant. On Revilico, users typically:
* Upload or retrieve protein structures for each toxicity target
* Clean and standardize structures (remove irrelevant ligands, assign protonation)
* Define binding regions (known ligand sites or functional pockets)
For flexible or cryptic sites, **Protein Water MD** can later be used to generate alternate conformations.
This gives a set of docking-ready toxicity targets.
3. **Screen Leads Against Toxicity Targets**
Now test whether leads bind where they shouldn’t. This step focuses on risk identification before optimization procedures using generative chemistry.
Use **Static Docking** to:
* Rapidly screen each lead against the toxicity panel
* Identify strong or recurrent off-target binding signals
* Capture candidate binding poses for interpretation
**What you’re looking for**
* Unexpected strong predicted affinities
* Consistent binding across multiple toxicity targets
* Binding in functionally critical regions (e.g., pore, catalytic site)
This screen will give a target-by-compound off-target binding matrix that can be utilized to create a toxicity profile for your lead sets.
4. **Analyze Interaction Patterns and Mechanisms**
Not all binding equals toxicity, interpretation matters. Use **Pharmacophore Analysis** to:
* Identify interaction motifs driving off-target binding
* Compare these motifs to known toxicophores
* Distinguish promiscuous binders from target-specific interactions
This step answers:
* *Is toxicity risk structural or incidental?*
* *Which parts of the molecule are responsible?*
What results is a mechanistic hypothesis linking structure to toxicity risk.
5. **Validate High-Risk Interactions with Dynamics (Optional)**
For leads showing concerning signals, use **Protein-Ligand MD** to:
* Test whether off-target binding is stable over time
* Identify transient vs persistent interactions
* Eliminate docking false positives using extended dynamic time scales and different solvent conditions
This step is especially valuable for:
* Ion channels
* Flexible receptors
* Borderline docking results
This step will output a dynamic confirmation (or dismissal) of off-target binding risk, allowing for a more robust mechanistic understanding of off-target engagements.\
This step is especially important for kinase targets which share homology with several other kinases biologically.
6. **Predict System-Level Toxicity Endpoints**
Complement target-based screening with data-driven prediction. Use Revilico’s **ADMET-AI Engine** to:
* Predict toxicity endpoints (hERG, DILI, AMES, LD50, etc.)
* Assess metabolic inhibition and clearance risks
* Surface confidence-scored alerts and red flags
* These models are trained on a wide expanse of experimental data to help get a picture of toxicity across different parameters
This step provides a global safety perspective beyond individual targets, and a consolidated toxicity risk profile per lead.
## **Results**
* Early identification of potential toxicity liabilities
* Mechanistic insight into off-target interactions
* Reduced risk of late-stage failure
* Clear guidance on which leads to advance, redesign, or deprioritize
* The mechanistic insight gained from these engines will allow for better optimization of downstream compound sets especially when it comes to toxicity
## **Now What?** *I’ve identified toxicity risks, but what’s next?*
Typical next steps include:
* Structural redesign to remove toxicophores
* Scaffold hopping or motif preservation workflows
* Focused experimental safety assays
* Re-screening optimized compounds against the toxicity panel
## **Why Revilico?**
Revilico enables proactive toxicity assessment by integrating:
* Target-based virtual screening for off-target risk
* Interaction-level analysis for mechanistic clarity
* AI-driven toxicity prediction for system-wide safety insight
* Seamless iteration into redesign and optimization workflows
This allows teams to fail fast, learn early, and design safer molecules before toxicity becomes an expensive problem.
# Quantum Chemistry Characterization
Source: https://docs.revilico.bio/solutions/lead-optimization/Deeply-Characterize-Lead-Molecules
Deeply Characterize Lead Molecules Using Quantum Chemistry and Property Analysis
## **The Problem You Are Trying to Solve**
*“I have lead compounds, and I want a deeper understanding of their molecular properties, such as metabolic stability, conformational behavior, electronic structure, and physicochemical properties, to guide confident lead optimization decisions.”*
At the lead optimization stage, activity alone is no longer sufficient. Teams often struggle with:
* Unclear metabolic liabilities or reactive hot spots
* Conformational flexibility that undermines binding or selectivity
* Poor solubility or formulation risk discovered too late
* Limited intuition for *why* small changes dramatically affect behavior
This workflow is designed to convert leads from black boxes into well-understood molecular systems, enabling rational optimization and risk reduction before expensive synthesis or in vivo work.
## **Solution**
This workflow leverages Revilico’s quantum chemistry and property prediction engines to provide a multi-layered, physics-informed understanding of lead molecules. The primary analysis chain is: Conformer Search → Geometry Minimization & Thermochemistry → Molecular Orbital Analysis → ADMET-AI & Solubility → Retrosynthesis Feasibility.
Together, these steps explain:
* *What shapes the molecule adopts, and the composition of these shapes in different contexts*
* *How stable and reactive it is electronically, and what is the most likely geometry it takes on*
* *How it is likely to behave in biological, metabolic, and formulation contexts*
* *Whether it is practical to make and scale*
Integration with binding chemistry engines (Docking, MD, FEP) is possible before or after molecular-level risks or hypotheses are identified.
## **What Data Do I Need to Provide?**
Required
* Lead compound structures (SMILES, CSV, or SDF)
Recommended
* A small set of close analogs (for comparison and trend analysis)
* Any known experimental liabilities (clearance, solubility, toxicity flags)
Optional
* Target binding poses (to contextualize conformational or electronic findings)
* Proposed chemical modifications under consideration
## **Workflow**
1. **Enumerate Realistic Molecular Conformations**
Begin by understanding how the molecule behaves in 3D space. Use **Conformer Search** to:
* Generate a diverse ensemble of low-energy conformers
* Quantify flexibility, shape diversity, and Boltzmann populations
* Identify dominant conformations versus rare, strained geometries
This step answers:
* *Is the molecule rigid or highly flexible?*
* *Does it naturally adopt binding-compatible shapes?*
* *Are there high-energy conformations likely to cause strain or instability?*
This step gives a curated conformational ensemble with energies and population weights.
2. **Optimize Geometry and Assess Thermodynamic Stability**
Next, refine molecular structures at quantum-mechanical resolution. Use **Geometry Minimization and Thermochemistry** to:
* Optimize geometries using ab-initio / DFT methods along with Neural Network Potentials and Machine Learning Interatomic Potential models
* Compute electronic energy, enthalpy, and Gibbs free energy
* Detect unstable geometries or imaginary vibrational modes
This step helps determine:
* Intrinsic molecular stability
* Relative favorability of conformers or analogs
* Whether structural modifications introduce energetic penalties
This results in optimized 3D geometries and thermodynamic state functions.
3. **Analyze Electronic Structure and Reactivity**
### To understand reactivity, metabolism, and potential liabilities, examine electronic properties. Use **Molecular Orbital (HOMO-LUMO) Analysis** to:
* Compute HOMO/LUMO energies and gaps
* Visualize orbital distributions
* Assess electronic hotspots associated with reactivity or metabolism
This step informs:
* Likelihood of oxidative metabolism
* Potential covalent or reactive behavior
* Stability versus promiscuous reactivity
This results in orbital energies, gap analysis, and 3D electronic visualizations.
4. **Evaluate Developability and Safety Risk**
Now translate molecular features into biological risk. Use **ADMET-AI** to predict:
* Absorption and permeability
* Metabolic clearance and CYP interactions
* Toxicity risks (hERG, DILI, AMES, etc.)
* Key physicochemical properties (LogP, TPSA, RO5 compliance)
This step provides early answers to:
* *Why might this lead fail in vivo?*
* *Which properties most urgently need improvement?*
* *Are risks structural or tunable?*
This engine gives you confidence-scored ADMET profiles with interpretable alerts.
5. **Assess Solubility and Formulation Risk**
### Before committing to synthesis or scale-up, assess solubility behavior. Use **Compound Solubility** to:
* Predict solubility across solvents and temperatures
* Identify formulation-friendly conditions
* Flag solubility-driven bioavailability risk
This step is especially important when:
* Leads are lipophilic or aromatic-heavy
* Crystallization or formulation has been challenging historically
This step provides solubility curves and solvent-specific recommendations.
6. **Evaluate Synthetic Feasibility**
Finally, ensure your lead is practical to make. Use **Retrosynthesis** to:
* Propose ranked synthetic routes
* Identify starting material availability
* Highlight steps or motifs driving synthetic complexity
This step answers:
* *Can this lead realistically be made and optimized?*
* *Which modifications are synthetically tractable?*
This engine will output ranked retrosynthetic pathways with feasibility scores.
## **Results**
* A physics-grounded understanding of lead conformations, stability, and electronics
* Early identification of metabolic, solubility, and toxicity risks
* Clear structure–property relationships guiding next optimization steps
* Confidence that selected leads are not only active, but viable
## **Now What?** *I understand my lead’s molecular behavior, but what’s next?*
Typical next steps include:
* Targeted structural modifications informed by Revilico’s Quantum and ADMET insights
* Multi-parameter optimization using generative chemistry
* Binding refinement using docking, MD, or FEP
* Selecting candidates for synthesis and experimental validation
## **Why Revilico?**
Revilico enables deep lead characterization by integrating:
* Quantum chemistry for first-principles understanding
* AI property prediction for scalable risk assessment
* Conformational analysis for realistic 3D behavior
* Synthesis-aware planning to keep design grounded in reality
This allows teams to move into late-stage optimization with clarity, confidence, and fewer surprises.
# ADME Multi-Parameter Optimization
Source: https://docs.revilico.bio/solutions/lead-optimization/Multi-Parameter-Optimization-for-Developability
Multi-Parameter Optimization for Developability: Solubility, ADME, and Toxicity
## **The Problem You Are Trying to Solve**
*“I have a set of lead compounds, and I want to optimize them across multiple developability dimensions, like solubility, ADME, toxicity, and related properties, without losing on-target potency.”*
At the lead optimization stage, teams often face tradeoffs:
* A potent binder may be insoluble or poorly permeable
* A promising series may carry CYP liabilities, hERG risk, or DILI flags
* Small structural changes can meaningfully shift clearance, bioavailability, or safety
Multi-parameter optimization (MPO) is about balancing the full profile, not maximizing a single metric.
## **Solution**
This workflow uses a primary MPO loop that alternates between:
1. Designing new analogs with the right constraints
2. Scoring them across developability endpoints
3. Filtering for synthesize-ability and feasibility
4. Repeating until a balanced set emerges
The primary MPO chain on Revilico is: Baseline Profiling → Property Screening (Solubility + ADMET-AI) → Multi-Objective Molecular Optimization → Synthesis Feasibility (Retrosynthesis) → Iterate.
Binding engines can be integrated at any stage to ensure potency and selectivity are not sacrificed while optimizing developability.
## **What Data Do I Need to Provide?**
Required
* Lead structures (SMILES / CSV)
Recommended
* Your desired optimization targets (examples: “increase solubility, reduce CYP3A4 inhibition risk, maintain MW \< 500”)
* Optional hard constraints (avoid specific motifs, preserve pharmacophore features, keep series scaffold)
Optional
* Experimental solubility/clearance/tox data (if available)
* On-target structure/assay signals for potency anchoring
## **Workflow**
1. **Generate New Analogs Using Multi-Objective Molecular Optimization**
Design new molecules that improve multiple endpoints simultaneously. Use **Generative Chemistry → Molecular Optimization** to:
* Generate analogs from the existing leads (conservative or exploratory)
* Ensure that there is proper activity of your ligands to the target of interest during this process
* Encode MPO goals as a multi-component scoring function (e.g., solubility ↑, CYP risk ↓, hERG risk ↓, MW/LogP constraints, etc.)
* Apply diversity controls so you don’t get 1,000 near-duplicates
This step is the core MPO engine: sample → score → update → repeat, producing candidates increasingly aligned to your objective profile, resulting in a versioned library of optimized analogs ranked by multi-objective score
2. **Establish the MPO Baseline**
After generating new molecules, define the success criteria for the lead series. On Revilico, users typically:
* Run **ADMET-AI Analysis** to surface the major liabilities (toxicity flags, CYP inhibition/substrate risks, clearance risk, permeability proxies)
* Run **Compound Solubility** to understand baseline solubility trends across solvents and temperature
* You can run the **Ligand Membrane MD** Engine that will allow you to analyze permeability of the compound to different seeded cell membranes
* (Optional) Run **Conformer Search** or other quantum methods to understand flexibility and shape drivers that may impact permeability and exposure
This baseline informs:
* Which liabilities matter most
* Which constraints must remain fixed
* How aggressive optimization should be
This step provides a clear MPO objective (prioritized endpoints + acceptable ranges).
3. **Rapid Property Screening and Downselection**
Run fast, high-throughput filters to avoid spending time optimizing the wrong chemistry. Use:
* **ADMET-AI Analysis** to identify red flags early (toxicity, CYP panels, hERG/DILI risk, clearance risk, permeability)
* **Compound Solubility** to identify which leads have a solubility ceiling that must be addressed by design
At this stage, you are not looking for perfection, you are identifying:
* which leads are salvageable
* which liabilities are dominant
* which properties move together vs trade off
This gives a ranked lead shortlist + a prioritized “liability map” per compound.
4. **Interpret and Filter Results**
After the optimization run, identify candidates that are *balanced*, not just “good at one metric.” Common selection heuristics:
* Prefer compounds with consistent ADMET improvements (not one extreme gain paired with new liabilities)
* Avoid candidates that only score well due to compensation effects
* Maintain chemical diversity among finalists to reduce risk of series-level blind spots
Use:
* **ADMET-AI Analysis** (re-run on the new library) for confirmation
* **Compound Solubility** (re-run for top candidates) to validate solubility gains
This produces a short list of MPO-balanced candidates for synthesis planning
5. **Check Conformations, Electronic Risk, and Physical Plausibility (Optional)**
For finalists, deepen confidence before synthesis. Use:
* **Conformer Search** to ensure the compounds can adopt relevant 3D shapes without extreme strain
* **Geometry Minimization and Thermochemistry** to confirm stable geometries and rule out unstable/high-energy structures
* **Molecular Orbital Analysis (HOMO–LUMO)** for electronic reactivity signals (useful for flagging unstable/reactive chemistry)
This step supports “sanity checks” that catch hidden risks, giving physics-backed confirmation that finalists are structurally and electronically reasonable.
6. **Ensure Synthesizability with Retrosynthesis**
A candidate that cannot be made is not a lead. Use **Retrosynthesis** to:
* Propose ranked synthetic routes
* Flag candidates with unrealistic or costly synthesis pathways
* Prioritize compounds that are both *better* and *buildable*
This engine results in a buildable shortlist with actionable synthetic plans.
## **Results**
* A refined set of candidates improved across solubility, ADME, and toxicity dimensions
* Explicit tradeoff awareness (what improved, what worsened, and why)
* Higher confidence in which compounds are worth synthesis and assays
* A repeatable MPO loop you can run iteratively as new data arrives
## **Integration with Other Engines (Optional)**
If potency/selectivity must be enforced within the MPO loop, users can integrate:
* Docking / Virtual Screening to maintain binding hypotheses while improving developability
* Protein–Ligand MD to validate binding stability for optimized candidates
* ABFE/RBFE for high-confidence ranking among close analogs
This allows MPO to be developability-first without losing efficacy.
## **Why Revilico?**
Revilico enables MPO as a single connected loop:
* Generative design (**Molecular Optimization**) grounded in explicit scoring
* Fast developability screens (**ADMET-AI**, **Solubility**) to guide iteration
* Physical plausibility checks (**Conformer Search**, **QM engines**) for confidence
* Buildability validation (**Retrosynthesis**) to ensure real-world feasibility
This produces candidates that aren’t just more potent, but actually more likely to succeed in development.
# Design Around Patents
Source: https://docs.revilico.bio/solutions/lead-optimization/Non-Infringing-Chemical-Matter-with-Comparable-Performance
Design Novel, Non-Infringing Chemical Matter with Comparable Performance
## **The Problem You Are Trying to Solve**
*“I have a set of lead compounds that infringe on existing patents. I need to design novel chemical matter that preserves favorable performance while avoiding infringement.”*
At the late lead optimization stage, patent risk becomes a gating constraint. Even highly effective compounds may be unusable if they:
* Fall within protected scaffolds or claim language
* Are obvious analogs under doctrine-of-equivalents reasoning
* Depend on protected substitutions, linkers, or core motifs
This workflow helps you systematically escape patent space while maintaining on-target performance, using physics-based validation and AI-guided redesign.
## **Solution**
This workflow centers on controlled chemical novelty:
* Preserve *functional performance* (binding mode, interactions, properties)
* Alter *chemical identity* (scaffold, connectivity, electronics, shape) enough to establish novelty
The primary redesign chain is: Patent-Aware Deconstruction → Motif Preservation → Novel Chemistry Generation → Performance Re-validation → Developability & Synthesis Checks.
Binding, quantum chemistry, and AI engines work together to ensure replacements are both non-infringing and credible leads.
## **What Data Do I Need to Provide?**
Required
* Lead compounds as SMILES (the infringing set)
* Target protein structure (experimental or modeled)
Recommended
* Knowledge of *why* the compound performs well (binding mode, key residues)
* Known patent boundaries (claimed scaffolds, functional groups, substitution rules)
Optional
* Internal SAR or experimental benchmarks
* Property constraints (ADMET, solubility, toxicity ceilings)
## **Workflow**
1. **Deconstruct the Infringing Leads**
Before changing chemistry, identify what *cannot* be lost. Users typically:
* Analyze docking poses or experimental complexes
* Identify key interaction motifs (H-bond donors/acceptors, hydrophobic anchors, π-stacking regions)
* Separate *performance-critical features* from *chemically incidental features*
* By understanding structure activity relationships and critical motifs that drive favorable binding or other properties, you can ensure preservation of the key scaffolds on the molecules during expansion efforts.
Primary engines used:
* **Docking / Ensemble Docking** to visualize conserved binding modes
* **Protein-Ligand MD** to confirm which interactions are stable vs. incidental
* **Pharmacophore Analysis** to abstract interactions into claim-agnostic features
* **QSAR Modeling** to get a better picture of the chemical space and which substructures are driving favorable properties of your leads
This step answers:
* *What makes this molecule work?*
* *Which elements must be preserved in function, not structure?*
This step will give you a functional interaction map and motif definition, independent of patented chemistry. We would recommend taking this key scaffold and running an independent search on the patent databases like sureChemBL.
2. **Map the Patent Risk Surface**
With functional motifs defined, shift focus to *where you cannot go*. Users typically:
* Identify scaffold cores, linkers, or substituent patterns that overlap claims
* Flag chemical regions that are too close to existing exemplified compounds
* Decide whether novelty should come from:
* Scaffold hopping
* Conformational reshaping
* Complete reconstruction of the molecule
This step informs how aggressive the redesign must be, providing clear “no-go” regions and acceptable novelty strategies.\
You can begin this process by throwing your compound into SureChemBL, searching for infringing patents, and diving into the details to know the breadth of patents for your critical scaffolds.
3. **Generate Novel Chemical Matter**
Use Generative Chemistry engines to design new molecules that:
* Preserve pharmacophore features
* Replace patented scaffolds or connectivity
* Explore chemically distinct regions of space
Primary engines:
* **De Novo Library Generation** for scaffold-level novelty
* **Molecular Optimization** for controlled, constraint-aware redesign, geared towards property optimizations
* **Custom Model Training** (optional) to bias generation toward internal success criteria
Generation is typically constrained by:
* Required interaction features (from Step 1)
* Property and synthesizability bounds
This will provide novel, patent-diverse compound libraries aligned to the same biological objective. You will need to eventually re-screen your generated set of molecules into SureChemBL for your high priority lead molecules.
4. **Re-Validate Performance Against the Target**
Novelty alone is not enough; new compounds must still work. Primary engines:
* **Docking → Flexible / Ensemble Docking** to confirm binding feasibility
* **Protein–Ligand MD** to test pose stability and interaction persistence
* **Free Energy Perturbation (RBFE / ABFE)** to quantify performance relative to the original lead
This step answers:
* *Do these new molecules bind the same way (functionally)?*
* *Have we preserved or improved potency?*
This will result in a ranked set of non-infringing candidates with validated on-target engagement.
5. **Evaluate Novelty-Driven Property Risk**
Structural novelty can introduce new liabilities. Users typically assess:
* **ADMET-AI** for toxicity, metabolism, clearance, and safety flags
* **Compound Solubility** to avoid formulation regressions
**Other Quantum Properties** to ensure that the structure of the molecule has the right geometries, electronic properties, and metabolic profiles.\
This step ensures novelty did not come at the cost of developability, giving a refined shortlist of viable, novel leads.
6. **Confirm Buildability and Freedom-to-Operate (Optional)**
Before committing, ensure redesigned compounds are practical. Optional next steps include:
* **Retrosynthesis** to verify synthetic accessibility
* **SureChemBL** patent database to search the novel structures for infringing patents based on tanimoto similarities and other criteria
* You can utilize all the other engines we offer to re-screen your new molecules to maintain favorable property profiles.
This will produce synthesis-ready, patent-aware lead candidates.
## **Results**
* Chemically distinct leads that preserve functional performance
* Reduced patent infringement risk through scaffold and topology changes
* Quantitative confidence that performance has been retained or improved
* A clear audit trail from patented lead → novel candidate
**Now what?**
* After receiving your lead set, make sure to review the new compounds you’d like to take into synthesis into the **SureChemBL** Engine which will allow you to confirm patentability of your molecules
* You can re-screen all of your molecules using the rest of the Revilico Engines for properties you are interested in.
## **Integration with Other Engines (Optional)**
This workflow integrates naturally with:
* Multi-parameter optimization pipelines (potency, ADMET, solubility)
* Toxicity and off-target screening workflows
* Experimental planning via prioritized synthesis routes
## **Why Revilico?**
Revilico uniquely enables patent-aware lead redesign by combining:
* Generative chemistry with hard biological and chemical constraints
* Physics-based validation (MD, FEP) to protect performance
* Quantum and property engines to manage novelty risk
* AI-assisted interpretation to accelerate decision-making
This allows teams to move beyond “cosmetic analogs” and confidently develop truly novel chemical matter that stands up scientifically, commercially, and legally.
# Planning Retrosynthesis
Source: https://docs.revilico.bio/solutions/lead-optimization/Plan-Synthetic-Routes-and-Pathways-for-Lead-Compounds
Plan Synthetic Routes and Pathways for Lead Compounds
## **The Problem You Are Trying to Solve**
*“I have a set of lead compounds, and I want to plan practical synthetic routes and pathways so I can prioritize what to make next.”*
At the lead optimization stage, the best molecule on paper is only valuable if it is buildable. Synthetic planning is often constrained by:
* Limited starting material availability
* Route length, yield, and step complexity
* Risky transformations or fragile intermediates
* Cost, cycle time, and scalability considerations
This workflow helps you move from candidate SMILES → actionable synthetic routes, with clear prioritization signals across your lead set.
## **Solution**
This workflow uses Revilico’s **Retrosynthesis engine** as the core route-planning layer, supported by optional feasibility and risk checks from other chemistry engines. The primary synthetic planning chain is: Lead Set Preparation → Retrosynthesis Route Generation → Route Ranking & Comparison → Starting Material & Risk Review → Export + Iterate.
Binding, property, and AI engines can be integrated to ensure you prioritize the compounds that are both valuable and makeable.
## **What Data Do I Need to Provide?**
Required
* Lead structures as SMILES (CSV upload or manual input)
Recommended
* Optional compound identifiers (name / series / project tag)
* Any “hard constraints” from your chem team (must avoid certain reagents, protect certain groups, limit step count, etc.)
Optional
* Known preferred intermediates or supplier catalogs (if your org has them)
* Target number of steps / cost ceilings / timeline constraints
## **Workflow**
1. **Prepare and Organize Your Lead Set**
Start by ensuring your lead list is clean and trackable. On Revilico, users typically:
* Upload a CSV of lead SMILES (plus optional IDs/names)
* Confirm structures are valid and standardized
* Group compounds into series (if relevant) so you can compare routes across analog families
This will give you a versioned lead set ready for route planning.
2. **Generate Retrosynthetic Pathways**
Use **Retrosynthesis** to propose diverse synthetic routes for each lead. This engine will:
* Expand multiple disconnection strategies per molecule
* Produce stepwise pathways with intermediates and reaction class labels
* Rank routes by a route-quality score (route plausibility, efficiency, starting material reasonableness)
This step answers:
* *How would I make this?*
* *How many routes exist?*
* *Which ones look most realistic?*
This will produce a ranked set of retrosynthetic pathways per lead.
3. **Compare and Prioritize Routes Across Leads**
Now shift from route generation to decision-making. Users typically review:
* Step count (shorter is usually faster and lower risk)
* Route diversity (multiple independent options reduces project fragility)
* Intermediate complexity (risk of bottlenecks)
* Starting material practicality (availability and cost proxies)
* Convergence opportunities (shared intermediates across a series)
* At this point, all of the data can be sent to your chemistry team for utilizing these newly generated hypotheses as a baseline for getting these molecules synthesized
This step is where you decide which molecules are:
* Ready to make now
* Worth minor redesign to reduce synthesis complexity
* Not currently practical relative to alternatives
This will give you a prioritized list of leads based on synthetic feasibility and route quality.
4. **Sanity-Check Molecular Feasibility and Stability (Optional)**
For routes that look promising but uncertain, validate that candidates and intermediates are physically reasonable.\
Optional supporting engines:
* **Geometry Minimization and Thermochemistry** to sanity-check stable geometries and identify strained or unstable candidates
* **Molecular Orbital Analysis (HOMO–LUMO)** to flag potentially reactive or unstable electronic profiles (helpful for identifying “looks good, but might be chemically problematic” cases)
* **Conformer Search** to highlight extreme flexibility or conformational strain that could complicate synthesis or isolation
This will give risk flags and confidence boosts on route feasibility.
5. **Export Routes and Create a Make-List**
Once routes are selected, generate outputs that enable execution:
* Route summaries per compound
* Stepwise reaction outlines and intermediates
* A consolidated “Make Next” list for your synthesis team
This step is designed to reduce handoff friction from computational planning → wet lab execution with a synthesis-ready route package and execution shortlist.
**Now what?**
* With the data on hand for planning synthesis, your chemists can now move forward with getting the molecules synthesized.
* If you are using this engine as a secondary screen to any other engine, you can utilize the results as a sanity check of the compounds to ensure that the compounds made with generative chemistry are feasible to move forward with.
## **Integration with Other Engines (Optional)**
Synthetic planning rarely happens in isolation. Revilico supports tight integration with:
* ADMET-AI + Solubility to avoid planning routes for compounds likely to fail developability
* Docking / MD / FEP to ensure synthesis effort is directed toward leads with strong on-target justification
* Generative Chemistry to redesign hard-to-make leads into more synthesizable analogs while preserving activity motifs
## **Why Revilico?**
Revilico makes synthetic planning actionable by combining:
* A dedicated retrosynthesis engine for route generation and ranking
* Optional physics/QM checks for stability and feasibility confidence
* AI-assisted interpretation to speed prioritization and iteration
* Integrations with design, screening, and optimization workflows so you can plan synthesis for the *right* compounds, not just the most interesting ones on paper.
# On and Off Target Effects
Source: https://docs.revilico.bio/solutions/lead-optimization/Understand-On-Target-and-Off-Target
Understand On-Target and Off-Target Biological Pathway Engagement for Lead Compounds
## **The Problem You Are Trying to Solve**
*“I have a set of lead compounds, and I want to understand both their direct on-target effects and their indirect off-target effects, including how they engage broader biological pathways and cellular programs.”*
At the lead optimization stage, it is no longer sufficient to know *that* a compound binds a target. You need to understand:
* Whether the compound robustly engages the intended target in a biologically meaningful way
* Which secondary proteins or pathways may be perturbed downstream
* Whether observed phenotypes arise from on-target mechanism, off-target liabilities, or network-level effects
* How molecular binding events translate into cellular state changes
This workflow bridges molecular-scale binding analysis with systems-level biological response modeling, helping teams move from “binds the target” to “behaves as expected in biology.”
## **Solution**
This workflow integrates binding chemistry engines with transcriptomic pathway analysis to connect molecular engagement with biological consequence. The primary analysis chain is: On-Target Binding Validation → Off-Target Binding Survey → Transcriptomic Response Analysis → Pathway & Network Interpretation.
This allows users to:
* Confirm intended target engagement
* Identify plausible off-target interactions
* Observe how cells respond to compound treatment
* Resolve whether effects are direct, indirect, or compensatory
Other engines (MD, FEP, quantum chemistry, generative chemistry) can be integrated as follow-ups once mechanisms are identified.
## **What Data Do I Need to Provide?**
Required
* Lead compound structures (SMILES or CSV)
* Primary target protein structure or sequence
* Cellular transcriptomic data (e.g., scRNA-seq) from treated vs control conditions
Recommended
* Time-resolved transcriptomic data (multiple doses or timepoints)
* Known pathway annotations for the disease context
Optional
* Off-target protein panels
* Comparative compounds (tool compounds, inactive controls)
## **Workflow**
1. **Validate Direct On-Target Engagement**
Start by confirming that leads engage the intended target in a plausible and specific manner.\
Using **Docking** and (if applicable) **Boltz or BoltzGen Co-Folding**, users:
* Evaluate binding poses and interaction patterns
* Confirm engagement of known functional residues
* Compare multiple leads for consistency of on-target binding
This step establishes a mechanistic anchor for downstream interpretation, giving a validated on-target binding hypotheses, and key interaction residues and motifs.
2. **Survey Potential Off-Target Binding**
Next, assess whether leads may engage additional proteins that could drive indirect effects. Using **Virtual Screening** and **Docking** against a curated off-target panel, users can:
* Identify secondary binding candidates
* Flag proteins involved in signaling, metabolism, or stress responses
* Distinguish selective compounds from promiscuous binders
This step generates candidate off-target hypotheses, not conclusions, in the form of a ranked list of plausible off-target interactions per lead.
3. **Analyze Cellular Response with scRNA-seq**
Now connect molecular interactions to biological outcomes. Using **scRNA-seq Analysis**, users can:
* Perform differential gene expression (DEG) analysis between treated and control cells
* Identify transcriptional programs altered by compound exposure
* Resolve heterogeneity across cell populations
This step answers:
* *What changes when the compound is applied?*
* *Which cell states are most affected?*
This results in differential expression profiles across cell types and conditions.
4. **Map Effects to Pathways and Gene Regulatory Networks**
Translate gene-level changes into mechanistic understanding. Using **Automated Target ID and Gene Regulatory Network Analysis**, users can:
* Identify upstream regulators driving observed transcriptional changes
* Map altered genes to known signaling pathways
* Predict cascade effects downstream of target engagement
This helps distinguish:
* Direct on-target pathway modulation
* Indirect off-target or compensatory responses
This step gives a pathway-level interpretation of compound effects with ranked regulators and network hubs.
5. **Model Temporal and Dose-Dependent Effects (Optional)**
If time-series data is available, deepen the analysis. Using **Temporal Omics Analysis**, users can:
* Model how gene expression evolves after compound exposure
* Identify early vs late response programs
* Separate primary effects from downstream adaptation
This step is especially valuable for:
* Chronic dosing scenarios
* Pathway rewiring and resistance studies
This produces dynamic models of pathway engagement over time.
6. **Integrate and Interpret with Revilico Agent**
Finally, synthesize results across molecular and cellular layers. Using **Revilico Agent**, users can:
* Ask mechanistic “why” questions across datasets
* Compare on-target binding hypotheses with observed pathway changes
* Generate testable biological hypotheses for follow-up experiments
This gives cohesive mechanistic narratives linking chemistry to biology.
## **Results**
* Clear differentiation between on-target and off-target effects
* Mechanistic insight into pathway-level engagement
* Reduced ambiguity in phenotypic interpretation
* Stronger confidence in lead progression or redesign decisions
## **Now What?** *I understand how my leads engage biology, but what’s next?*
Common next steps include:
* Refining chemistry to amplify desired pathways and suppress undesired ones
* Designing focused validation experiments
* Prioritizing leads with clean, interpretable mechanisms
* Integrating findings into toxicity or efficacy optimization workflows
## **Why Revilico?**
Revilico uniquely connects:
* Molecular binding engines for direct interaction analysis
* Single-cell transcriptomics for biological response resolution
* Network and temporal modeling for pathway-level insight
* AI-assisted reasoning to unify complex datasets
This enables teams to move beyond single-target thinking and design molecules with predictable, system-aware biological behavior.
# Centralized Experimental + Computational Data Management
Source: https://docs.revilico.bio/solutions/operational-efficiency/Centralized-Experimental-Computational-Data-Management
“I want my data to live in one place where it is visible, usable, and computationally actionable.”
### **The Problem**
Scientific data often becomes fragmented across multiple platforms: experimental assay results reside in ELN/LIMS systems, computational outputs are stored in separate simulation folders, and visualizations require external tools. As a result, data must be reformatted before being used for modeling or screening, and a lack of drag-and-drop interoperability exists between experiments and computation. This leads to slow insight generation, suffers reproducibility, computational screens disconnected from real experimental data, and teams struggling to maintain a single source of truth.
### **The Solution**
Revilico addresses this by centralizing experimental and computational workflows into a unified ecosystem. Key features include Drive for storing all raw and processed data, a Spreadsheet Editor for rapid visualization and sorting of assay data. The 2D and 3D Viewer provides interactive 3D structural visualization directly from stored files, The Central Hub monitors pipeline status, the Project Hub organizes programs, and drag-and-drop integration moves assay CSV files directly into Docking, QSAR, ADMET-AI, or Generative workflows. Furthermore, the Compound Review + Design Suite allows direct structural editing, and Revilico GPT/Agent & Interpreter generate summaries and quick analyses from stored files.
### **Outcome**
With this unified ecosystem, your experimental and computational data become: Centralized (one cohesive storage and access layer), Interoperable (instantly usable for modeling and screening), Visualizable (no exporting required), Collaborative (shareable with context preserved), and Computationally ready (directly deployable into pipelines). Instead of moving data between systems, your data lives inside the same ecosystem where you can analyze, model, and simulate.
# Cross Team Collaborations
Source: https://docs.revilico.bio/solutions/operational-efficiency/Cross-Team-Collaboration-Without-File-Chaos
Cross-Team Collaboration Without File Chaos
*“My chemists, computationalists, and biologists need to work together without breaking each other’s workflows.”*
### **The Problem**
* Cross-functional collaboration often causes several challenges, including the proliferation of duplicate files across shared drives, unclear ownership of datasets, and confusion regarding file versions. Furthermore, the process frequently involves the manual transfer of structures between different modeling and design tools, as well as fragmented communication across various platforms such as Slack, email, and local storage.
### **The Solution**
* Revilico unifies collaboration by offering several features: Shared Drives centralize project files which can be shared across the entire team; Team Chat enables context-aware discussion tied to pipelines so that your slack communication also exists within the platform (after connecting); Compound Review supports shared evaluation and annotation of molecular libraries so that chemist’s intuitions are directly integrated into the organizations workflows; our 2D and 3D design suite enables chemists to directly modify stored structures and visualize different outcomes easily.
### **Outcome**
Collaboration becomes context-preserved, file-consistent, and computationally integrated. This ensures that design, simulation, and interpretation all happen within the same environment.
# Pipeline Transparency & Execution Oversight
Source: https://docs.revilico.bio/solutions/operational-efficiency/Pipeline-Transparency-Execution-Oversight
“I need to know what’s running, what failed, and what’s ready, without asking everyone.”
### **The Problem**
In computational discovery programs, efficiency and strategic oversight are hampered by several issues. Pipelines often run asynchronously across disparate engines, and team members frequently launch jobs without mutual visibility. This lack of coordination means failed jobs can go unnoticed, and without a unified view of computational progress, leadership struggles to assess project momentum. The overall result is the creation of delays, duplicated work, and a significant loss of strategic clarity.
### **The Solution**
Revilico addresses these transparency issues by providing centralized execution management. The system uses a Central Hub to display real-time pipeline status (such as Queued, Running, Completed, or Failed) and a Project Hub to organize pipelines by indication, target, or campaign. Pipeline Sharing allows for visibility across the team without file duplication. To maintain workflow integrity, the Revilico Agent can notify users of failed jobs or stalled workflows over slack or email, and Revilico Agent/GPT summarizes pipeline outputs for rapid reporting.
### **Outcome**
The ultimate outcome is that computational work becomes highly Visible, Auditable, and Strategically trackable. Instead of team members asking reactive questions like “Who ran this?” or “Is that done yet?”, the entire team operates with shared situational awareness and a clear understanding of the project's progress.
# Reproducibility, Traceability, Version Controls
Source: https://docs.revilico.bio/solutions/operational-efficiency/Reproducibility-Version-Control
Reproducibility & Version Control
*“I need to reproduce last quarter’s results without guessing what settings were used.”*
### **The Problem**
Scientific reproducibility breaks down when parameter configurations are undocumented, files are overwritten or renamed inconsistently, different team members use slightly different workflows, and results cannot be traced back to their input assumptions. This collectively makes validation, regulatory documentation, and scientific defense difficult.
### **The Solution**
Revilico embeds reproducibility into its operations by ensuring that pipelines automatically log your input parameters, force fields, and configuration settings, and they are all tied to your command center queues. Versioned outputs are stored in Revilico Drive, the Project Hub preserves all of the screens run within an organization on a per-project basis, and Pipeline Sharing keeps full metadata intact when transferring results within your team. Furthermore, the Revilico Interpreter allows quick inspection of parameter sets, and documentation pages per engine to clarify default versus user-modified settings.
### **Outcome**
The outcome is that every computational result has clear provenance, traceable inputs, and a reproducible configuration. Consequently, scientific conclusions are defensible and transparent.
# Structured Knowledge Transfer
Source: https://docs.revilico.bio/solutions/operational-efficiency/Structured-Knowledge-Transfer
“What happens if a key scientist leaves the project?”
### **The Problem**
When institutional knowledge is not structured, an organization faces significant issues, including undocumented workflow logic, lost parameter rationale, unclear pipeline configurations, and long onboarding periods for new hires. Also, if there is a team adaptation where a member leaves, and if things aren’t logged properly, it can become very difficult to manage the extensive data generated.
### **The Solution**
Revilico supports structured knowledge continuity by ensuring that pipelines remain saved with a full configuration history, documentation pages provide technical background and demos for those who have never utilized the engines before, the Project Hub maintains contextual grouping for different pipelines and projects, Revilico Agent can generate workflow explanations from pipeline metadata, and Shared Drives preserve institutional memory across the entire team. Furthermore, this system is designed for resilience against organizational changes, allowing data to be seamlessly transferred over and experiments to be logged in our notetaker without any impact on your existing data or workflows.
### **Outcome**
The scientific strategy is embedded directly into the platform, not locked in individuals, which makes onboarding faster and ensures continuity is automatic.
# Ligand Protein Multi-Plexing
Source: https://docs.revilico.bio/solutions/target-identification/Compound-Protein-Interaction-Discovery-at-Scale
Compound–Protein Interaction Discovery at Scale
## **The Problem You Are Trying to Solve**
*“I want to run a multiplex screen to understand how large compound libraries interact with large protein libraries, but the full experimental matrix is too expensive and time-intensive.”*
A true compound × protein screen can explode into millions of combinations, and even “lightweight” biochemical assays become impractical. Teams often need a way to:
* discover likely binders/interactors early
* triage the search space
* identify off-target risk and polypharmacology signals
* focus experimental validation on the smallest, highest-value subset
## **Solution**
This workflow uses Revilico’s binding chemistry engines to simulate compound–protein interaction likelihood at scale, then progressively increases rigor to produce a high-confidence list of protein interactors for any given compound (or compound interactors for any given protein). The primary refinement chain is: Virtual Screening / Docking → Pose & interaction QC → (Optional) MD stability → (Optional) FEP confirmation → Ranked interactor list.
You can run this workflow in either direction:
* **Compound-centric:** “What proteins does this compound likely bind?”
* **Target-centric:** “Which compounds bind this protein?”
* **Matrix mode:** “Score a compound library against a protein panel”
## **What Data Do I Need to Provide?**
Required
* Compound library as SMILES (CSV)
* Protein library as structures (PDB) or sequences (if you need structure generation first)
Recommended
* Known binding pockets or reference ligands (if available)
* Any existing experimental binding/activity data (even small) for calibration
Optional
* Desired selectivity constraints (e.g., avoid kinases, avoid hERG panel proteins)
* A protein “tox/off-target panel” list (for safety screening use cases)
## **Workflow**
1. **Assemble and Standardize Your Screening Inputs**
Before screening, ensure both libraries are in usable formats. Users typically:
* Upload compounds as a SMILES CSV
* Upload proteins as PDB files (one or many)
* Use platform utilities to convert/merge files as needed
This gives a clean compound library + a clean protein panel ready for screening.
2. **Establish the Screening Strategy**
At scale, you want a fast first pass with clear constraints. Users typically decide:
* What counts as an “interactor” (score threshold, pose confidence, binding mode plausibility)
* Whether to screen one compound vs many proteins, or many compounds vs a protein panel
* Whether known binding sites exist, or whether docking should search broader pockets using blind docking approaches
Primary engines:
* **Virtual Screening Engine (static, flexible, and ensemble docking)**
* **Boltz-Cofolding**
This provides a screening plan that prioritizes throughput and broad recall (catch candidates).
3. **Run High-Throughput Interaction Prediction**
This is the engine-driven multiplex substitute for wet-lab screening. Users run:
* **Virtual Screening or Boltz Co-Folding** (fast, high-throughput) across the protein panel and compound set
* If screening a *very large* compound library, start with **Static Docking** at scale, which allows for GPU optimized docking speed, and then down-select top candidates of interest.
This produces a large interaction score matrix (compound–protein pairs) with top-ranked candidates.
4. **Quality Control and Reduce False Positives**
Docking at scale is intentionally fast, so the next step is to remove obvious artifacts. Users typically:
* Filter out strained ligand poses or implausible geometries by looking at intramolecular energies or 3D conformational visualizations
* Require agreement across multiple poses (not a single outlier)
* Prioritize conserved interactions across related proteins (if relevant)
Primary engines and utilities:
* **Flexible Docking** (re-score a smaller batch, allow key residues to move)
* **Ensemble Docking** (if pocket flexibility is a concern)
* **Revilico Interpreter / RevilicoGPT** for automated interpretation of pose quality trends to extrapolate results of the data to experimental outcomes
This results in a refined list of likely interactors with improved precision.
5. **Confirm Stability in a Physical Environment (Optional)**
When you need higher confidence (or when docking is ambiguous), validate whether binding is stable over time. Users run **Protein Ligand MD** on the most important compound–protein pairs
**What you’re looking for:**
* Stable binding mode (no immediate dissociation / unrealistic drift) with meaningful breakdowns of energetic contributions
* Reasonable RMSD/RMSF behavior near the binding site
* Consistent key contacts over the trajectory
This will give a stability-validated subset of compound–protein interactions.
6. **Quantify Binding Strength for Final Confirmation (Optional)**
For the smallest shortlist where you want thermodynamic confidence: Users run:
* **MMPBSA/MMGBSA** binding energy scoring analysis across longer time scales
* **ABFE** to estimate absolute binding favorability for a compound–protein pair
* **RBFE** to compare close analogs (e.g., when ranking within a compound series)
This will produce a “gold” shortlist of interactions supported by physics-based free energy estimates, giving you more confidence when down selecting certain molecules for analysis in the lab.
## **Results**
* A **ranked, high-confidence list of protein interactors for any given compound** (or vice versa), including:
* predicted binding poses
* docking and re-scoring metrics
* optional MD stability evidence
* optional ABFE/RBFE thermodynamic support
This list is designed to replace a brute-force biochemical multiplex screen with a computational triage pipeline, reducing experimental burden to only the most valuable validations.
## **Integration with Other Engines (Optional)**
Depending on your downstream goal, this workflow can connect naturally to:
* **QSAR Modeling** (learn patterns from predicted/experimental activity)
* **Pharmacophore Analysis** (motif extraction for cross-target similarity)
* **Generative Chemistry** (design/selectivity tuning based on interaction profile)
* **ADMET-AI / Toxicity panels** (early developability triage)
* **Quantum Chemistry** (electronic property or reactivity checks for select pairs)
## **Why Revilico?**
Revilico supports multiplex interaction discovery because it combines:
* scale-first screening (Virtual Screening + Docking)
* precision refinement (Flexible/Ensemble Docking, MD)
* high-confidence confirmation (FEP)
* a unified interface for results inspection, interpretation, and iteration
* Ability to analyze these different factors across a wide variety of multi-plexed ligand protein interaction pairs at scale, simply and easily using simulations.
So instead of experimentally testing every combination, you computationally narrow the matrix to the few pairs that actually deserve bench time.
# Bioinformatic Analyses
Source: https://docs.revilico.bio/solutions/target-identification/Conducting-Multiomics-Analysis-on-Revilico-OS
Conducting Multiomics Analysis on Revilico OS
## The Problem You are Trying to Solve:
*“I want to conduct multiomics analysis with the data I have available to get a better understanding of my desired indication before diving into computational chemistry workflows.”*
Multiomics analysis is a powerful and holistic approach toward understanding the layers of a biological disease and its mechanisms, including the genomic, transcriptomic, and metabolomic fields. However, conducting multiomics analysis is often challenging, given the various layers and options for interpreting your data. Oftentimes, workflows for multiomics analysis can be scattered or require an individual with bioinformatics/data analytics skills to create custom scripts or pipelines to visualize and analyze results (pipelines that may vary between experiments). Having these models centralized into a unified platform enables users to not only conduct analysis with ease, but to do so through the various layers of this process.
## Solution
Here, we present the core features for multiomics analysis currently available on the Revilico platform. Our platform provides extensive workflows and features for transcriptomics and proteomics with additional upcoming features for genomics, metabolomics, and epigenomics in the near future. By using the currently available tools, our hope is that you will be able to have a better understanding of your data and your disease indication before you proceed through the next phases of your drug discovery pipeline.
### What Data Do I Need to Provide?
* Control and Experimental h5ad Files (scRNA-seq Analysis)
* CSV with Protein Sequences/Manually Added Protein Sequences (AlphaFold)
* Single Protein, Protein + Ligand, Multimer, or DNA with PTMs (OpenFold)
* PDB Files (Pocket Search)
## Workflow
1. ### scRNA-seq Analysis
To perform transcriptomics analysis by understanding the single-cell RNA sequencing profile of your data, start by uploading both your control and experimental sample data as h5ad files. The platform presents you with three possible analyses that you can select from: 1) DGE Analysis (identify significant changes in gene expression between your sample groups), 2) Automated Target ID (identify ideal targets from your cell line of interest), and 3) Temporal Omics Analysis (observe and analyze the gene expression in your data as a function of time for dynamic and temporal constraints). You can also select the top number of genes you want to extract to filter out the best possible results. Based on the chosen analyses, visualization plots and charts will be available in the results section for this specific tool. This will help you with identification of certain targets of interest using a transcriptomics based approach to move into later stage structure based drug design.
2. ### AlphaFold
For understanding the structure of your down-selected target/protein, utilize our data extraction AI agents to pull the amino acid sequence you will be utilizing for your protein model, or upload your protein sequence(s) of interest. Here, you have the option to fine-tune the parameters for your protein such as structure relaxation, templates, multiple sequence alignment, and pairing options. Protein structures will be available in the results section for this tool after running.
3. ### OpenFold
For a greater understanding of the structure of your protein using a different model with different intrinsic biases and performance, and including support for additional complexes, upload your queries based on the data structures accepted. An option is provided to tune multiple sequence alignment or to use templates to help gain more accuracy for your given complexes. 3D structures will be available in the results section for this tool after running.
4. ### Pocket Search
To identify top pockets of your protein of interest, upload your PDB files into Pocket Search Engine. Here, you have the opportunity to optimize sphere radii and chain selections for your use cases. You can also take this over to the MDPocket tab in this tool to combine your understanding of the available pockets with various molecular dynamic simulations available in our Binding Chemistry suite. Top binding pockets and their characteristics will be available in the results section for this tool after running. These engines help to find static pockets as well as more cryptic or transient pockets that can appear after longer dynamic time scales.
## Results
* DGE Analysis, Target ID, Temporal Omics (scRNA-seq Analysis)
* Protein Structures (AlphaFold)
* 3D Structures for Protein, Complexes, or DNA/RNA (OpenFold)
* Top Protein Pockets (Pocket Search)
By running the tools above and performing your analyses, you will be able to get a greater understanding of your disease indication and ideal targets you wish to proceed with. These tools elucidate the underlying biology behind your disease at the cellular pathway level and will enable you to generate structures based on the target or targets identified as potential therapeutic candidates.
### **Now what?**
I have a greater understanding of my disease indication and what target I want to go for!
* If you have molecules or a library of compounds you are interested in using → proceed through the various tools available in our Binding Chemistry or Quantum Chemistry suite to better understand your molecules' binding, docking, molecular dynamics, or quantum-level interactions.
* If you do not have molecules or a library of compounds → proceed to the De Novo Library Generation available in our Binding Chemistry suite to develop molecules for your newly selected target. You can also reference our pre-determined libraries that are on hand in liquid and power form.
## Why Revilico?
This Revilico workflow enables users to query the availability of experimental data before proceeding to AI structure prediction. All 3D structural hypotheses include transparent quality scoring, and can be seamlessly integrated with downstream discovery engines and workflows in a multi-modal way to help hedge all of your results against one another.
# Deciding Therapeutic Strategy
Source: https://docs.revilico.bio/solutions/target-identification/Determining-Optimal-Therapeutic-Strategy
Determining Optimal Therapeutic Strategy for Drugging a Molecular Target
**The Problem You are Trying to Solve:**\
*“I have identified a molecular target implicated in disease, and I want to determine the optimal therapeutic strategy/ mode of action (e.g., inhibition, activation, or allosteric modulation) to achieve the desired biological outcome”*
Traditionally, determine optimal therapeutic strategy required extensive experimental validation through costly and time-consuming assays: enzymatic activity screens to test inhibition, cell-based phenotypic assays to assess pathway modulation, co-immunoprecipitation experiments to validate protein-protein interactions, and often clinical trial data to reveal that the wrong modality was chosen after years of development. The disconnect between target biology, target structure, and cellular context meant therapeutic strategy selection was based on precedent rather than comprehensive mechanistic understanding. Now with AI-powered structural prediction, transcriptomic profiling, and computational binding analysis, we can integrate target structure, biological function, and systems level pathway effects to rationally predict optimal therapeutic approaches before committing to expensive experimental campaigns
**Solution**\
This workflow enables users to a) systematically evaluate all druggable sites on a target protein through structural pocket analysis and b) determine optimal therapeutic strategy and mode of action by integrating structural druggability, transcriptomic pathway analysis, and dynamic binding site characterization across Revilico’s multi-modal engine suite. We leverage complementary computational methods such as pocket identification for binding site discovery, scRNA-seq analysis for biological context and pathway validation, molecular dynamics for conformational flexibility assessment, and co-folding predictions for protein-protein interface analysis, to ensure therapeutic strategy selection is grounded in both structural feasibility and system-level biological understanding
**What Data Do I Need to Provide?**
* Protein sequence (used to generate the different possible protein structures with Alphafold)
**Workflow**
1. **Generate or Obtain Target Protein Structure**
Establish a high-confidence 3D structural model of your target protein to enable pocket identification and druggability assessment. We will first use **RevilicoGPT/Revilico Agent** with the following query: “I am evaluating AXL for Triple Negative Breast Cancer. Search PDB and retrieve any experimental structures, prioritizing high-resolution structures (\< 2.5 A) with bound ligands or in different functional states (apo, holo, active, inactive).” If there are no suitable experimental structures, we can run **Alphafold** or **Openfold** using our protein sequence, where we will prioritize structures with high confidence scores. Usually, our team likes to take co-crystal structures of known inhibitors with ideal modes of action to derive pocket coordinates and to understand ideal amino acids driving binding within key regions in the pocket.
2. **Identify Druggable Binding Sites**
We will run **Pocket Search Engine** on our protein structure to identify all cavities and rank by druggability score, volume, and physicochemical properties. Note pocket location, i.e. catalytic site pockets suggest orthosteric inhibition, surface pockets away from active sites suggest allosteric modulation, and float protein-protein interfaces suggest PPI disruption strategies.
For PPI targets, use **BoltzGen Co-Folding** to model the protein complex interface. Analyze interface area, binding energy, and hot-spot residues. Weak interfaces with shallow binding sites suggest PPI disruption is feasible; strong interfaces with deep binding grooves may require stabilization, alternative therapeutic modalities, or indirect modulation strategies.
3. **Analyze Biological Context Via Transcriptomics**
Use **Revilico Agent** to obtain relevant scRNA-seq datasets comparing disease vs control or perturbed vs baseline states for your target. Here is a sample query: “I am studying EGFR in Triple Negative Breast Cancer. Find scRNA-seq datasets in h5ad format comparing disease vs normal. The datasets must include EGFR and be from the cell line MDA-MB-231.” You will then run **DEG Analysis** and **Automated Target ID** from **scRNA-seq Analysis**. From the results we can extract the DEG volcano plot from DEG Analysis to see how differentially expressed our target gene is from control and experimental. From the Automated Target ID section we can see if downstream genes from our target are affected. We can also reference the phenotype bar graph to validate if our phenotype predictions of the target align with it. If we are seeing an over-activation of certain genes, we can use this information to design a mode of action that inhibits the translated protein function, therefore returning the system back to equilibrium.
4. **Assess Conformational Flexibility**
Run **Protein Water MD** simulations followed by **MDPocket** analysis to identify cryptic or transient products that only appear during protein motion. Cryptic pockets that open frequently represent opportunities for allosteric modulation that static analysis would miss. This extracts multiple protein conformations from the trajectory and shows which conformational states are thermally accessible. It will also reveal whether the protein is rigid or highly flexible.
5. **Validate Strategy with Proof-of-Concept Docking**
For orthosteric inhibition run **Static Docking** or **Flexible Docking** with known substrates, cofactors, or reference inhibitors against the active site pocket. Strong binding scores with chemically reasonable poses validate that the small molecule can effectively compete with natural substrates.
For allosteric modulation, dock small fragment libraries or known allosteric modulators into cryptic//allosteric pockets identified. Successful binding to these sites with stable poses suggest allosteric strategy is viable. Run **Protein-Ligand MD** on top poses to confirm allosteric pockets remain stable when occupied, and that they make conformational modifications to the entire protein structure (especially within the active site). This can be confirmed by looking at larger jumps in Root Mean Squared Fluctuatio (RMSF) within pocket residues.
For PPI Disruption, use **Boltzgen CoFolding** to model binding of peptide mimetics or small molecules at the protein-protein interface. Alternatively, dock fragment libraries targeting interface hot-spots. Validate using **Pharmacophore Analysis** that compounds can recapitalize key interface interactions and engage with key amino acids driving biological activity.
**Results**
* Ranked druggable pockets with druggability scores, volumes, and strategic classifications
* Transcriptomic response profiles
* Conformational flexibility assessment
* Proof-of-concept binding validation
* Recommended therapeutic strategy with structural and biological justification
Convergent evidence across structural druggability, biological phenotypes, and binding validation determines optimal strategy. For orthosteric inhibition: look for high scoring active site pockets, strong transcriptomic response showing target perturbation drives desired phenotypes, and validated substrate competitive binding. For allosteric modulation: prioritize cryptic pockets appearing >30% of MD simulation, moderate transcriptomic effects suggesting partial modulation suffices, and stable allosteric ligand binding in validation docking. For PPI disruption: weak interfaces, phenotypic predictions indicating complex formation drives pathology, and successful peptide/fragment binding at interface hot-spots. Discordant results suggest the target may require alternative modalities like protein degradation or may not be optimally druggable, prompting reconsideration of target selection or combination strategies.
**Now what?** After determining your mode of action and biophysical function of your system, you can move into iterative compound design.
* Utilizing all of the engines within the Revilico platform from binding chemistry to generative chemistry, you can run a variety of assessments on your compounds for further synthesis and testing.
**Why Revilico?**\
This workflow determines the optimal therapeutic strategy for a molecular target (e.g., inhibition, activation, or PPI disruption) by integrating several key data points. This involves assessing the target's structural druggability, confirming that perturbation drives the desired biological phenotype (transcriptomics), and validating a stable binding mode with compound studies.
# Generating 3D Structures
Source: https://docs.revilico.bio/solutions/target-identification/Generate-a-3D-Structural-Hypothesis-for-a-Molecular-Target
Generate a 3D Structural Hypothesis for a Molecular Target
**The Problem You are Trying to Solve:**\
*“I have a molecular target (protein) without an experimentally resolved structure, and I want to obtain a 3D structural hypothesis suitable for analysis and structure based drug design and discovery.”*
This problem can be traditionally difficult to solve, as many targets lack resolved structures, experimental structures may be incomplete or unavailable, and structure prediction must be interpreted with confidence to be valuable in downstream steps of drug discovery. Traditional methods include NMR, X-ray Crystallography, and Cryo-Electron Microscopy. Now with new developments in the AI space, it is becoming more prevalent to use engines to predict biological structures that are traditionally hard to resolve experimentally.
**Solution**\
This workflow enables users to a) determine whether a resolved structure exists or b) generate a high-confidence AI-predicted structural model using a variety of different protein folding and co-folding algorithms that are available on Revilico’s Operating System Platform. We also have a variety of engines for establishing co-crystal predicted structures of compounds bound to protein targets along with DNA, RNA, and Protein co-folding options to ensure you can properly represent your biological system computationally. Proper conformations of the ligand and confidences of the protein structure are calibrated using a model-hedging method to help negate singular model biases (i.e. Each algorithm has distinct training data and inductive assumptions, so by comparing and averaging across them, we reduce the resulting errors and increase confidence in our generated structures).
**What Data Do I Need to Provide?**
* Protein name or identifier (Required if you are searching databases for representative structures)
* Protein amino acid sequence (Required to use protein folding algorithms)
* Known ligands, DNA/RNA sequences, or binding partners (Optional, but required for algorithms that co-fold binding partners to protein targets)
* Desired binding pocket or conformation (Optional, but somewhat required for co-folding ligands)
* Other protein sequences (Optional, but required if you are looking at protein protein interactions)
**Workflow**
1. **Identify Existing Experimental Structures**
Determine whether a resolved structure already exists. To do this on Revilico, users can query structural databases (like the PDB or Uniprot) and evaluate resolution, coverage, and biological relevance with **RevilicoGPT** and **Revilico Agent.**
Sample Query: I am evaluating AXL for Triple Negative Breast Cancer and want a scope of all the publicly available protein structures I can use for structure based drug designs. Already known inhibitor co-crystal structures would be optimal.
RevilicoGPT and Revilico Agent will output experimental structure files (if available), and their corresponding coverage and quality annotations. If a suitable structure exists, users can proceed with downstream analysis using this obtained 3D structure. If no suitable structure exists, the user can proceed to AI structure prediction. Usually, we will look for target structures that already have well known co-crystal structures of inhibitors so that we can use it as a benchmark for our computational screening, and we take the structure into Pymol or ChimeraX to remove waters, the ligand, and ions to clean and prepare the structure for downstream docking and analysis.
Sometimes, you will also see in public databases that there are proteins that have full coverage or multimer configurations, but for proteins like kinases with a simplified kinase domain, we extract just the portion of the protein (with the active ATP binding site) and a ligand co-crystal (if available) to use as our baseline structure.
2. **Predict Structure from Sequence**
Generate a de novo 3D structural hypothesis. Taking a protein amino acid sequence as an input, the **AlphaFold, Boltz1/2,** and **OpenFold** engines will produce predicted 3D structure(s), per-residue confidence scores, and model provenance and metadata. To understand how to read out the different metrics, you can refer to the engine specific documentation here: AlphaFold, Boltz1/2, OpenFold.
It is important to note that outputs of AlphaFold, Boltz1/2, and OpenFold are labeled as hypotheses, not experimentally resolved structures. Proper analysis of the structure along with per residue PTM, pLDDT, and confidence scores should be assessed to help you gain confidence on your generated hypothesis.
3. **Ligand-DNA-RNA or Conformation-Specific Refinement (Optional)**
This step serves a specific case, when a specific binding pocket is known and a ligand or interacting protein is available. Users can input the protein structure or sequence, and ligand or binding partner into the Boltz or BoltzGen co-folding engines to produce protein-ligand or protein-protein complex structures with pocket-specific conformational hypotheses. Boltz2 co-folding is specifically geared towards generating ligand protein binding hypotheses and activity values. OpenFold is mainly utilized for protein protein interaction pairs, DNA-Protein, RNA-Protein, etc. biological pairs to get a more nuanced biological model for analysis.
**Results**
* Versioned 3D structural hypothesis
* Confidence and provenance metadata
* Structures ready for docking, simulation, or analysis
If you see that several models are in agreement on parameters and confidences, you can move forward with your computational campaign with greater trust in strong SBDD foundations. You should look for protein domain and region specific confidence scores to help validate assumptions on your target (i.e. flexible loops usually will have lower confidence, and structures that are more static/rigid should have higher predicted confidences)
**Now what?** I have my structure and want to run a SBDD campaign!
* **Still pending**: List of all the solutions that take in a protein 3D structure for SBDD
**Why Revilico?**\
This Revilico workflow enables users to query the availability of experimental data before proceeding to AI structure prediction. All 3D structural hypotheses include transparent quality scoring, and can be seamlessly integrated with downstream discovery engines and workflows in a multi-modal way to help hedge all of your results against one another.
# Target Identification
Source: https://docs.revilico.bio/solutions/target-identification/IdentifyPrimary-Molecular-Target-for-Phenotypic
Identify Primary Molecular Target for Phenotypic Hit via Reverse Docking
**The Problem You are Trying to Solve:**\
*“I have a drug with demonstrated efficacy but unknown mechanism of action, and I want to identify its primary molecular target(s) to enable structure-based optimization and rational drug design”*
Traditional target identification for phenotypic hits required labor-intensive biochemical approaches like affinity chromatography pull-down assays, proteomics screens, and genetic validation studies, each taking months and often yielding ambiguous results with multiple potential targets. Without knowing the binding site or target protein, rational optimization can be impossible, forcing researchers into costly trial and error medicinal chemistry campaigns that modify the drug blindly while monitoring phenotypic readouts. The disconnect between observable efficacy and molecular mechanism meant many promising drugs were abandoned or progressed to clinical trials without understanding their liabilities, leading to unexpected toxicities or failure to translate across disease contexts. Moreover, during the acquisition process of therapeutic candidates, knowing the target of interest is exceptionally important as it allows for a fall back to structure based drug design should anything go wrong during clinical trials or beyond.
**Solution**\
This workflow enables users to a) systematically identify primary molecular targets for phenotypic hits through reverse docking against disease-relevant protein libraries and b) validate target engagement through progressive refinement using Revico’s multi-tiered virtual screening and molecular dynamics engines. We leverage a cascade of increasingly rigorous computation methods, i.e. static docking for rapid screening, flexible docking for improved accuracy, ensemble docking for conformation sampling, protein-ligand MD for dynamic stability assessment, and free energy perturbation for quantitative affinity ranking at the highest level of accuracy. Proper target engagement validation is achieved through pharmacophore analysis that confirms reasonable binding modes and binding free energy calculations that distinguish true targets from docking artifacts.
**What Data Do I Need to Provide?**
* Known ligand or binding partner (required for docking, MD, FEP, and Pharmacophore analysis)
* Protein Library (Optional, we have libraries on hand, or we can generate this using **RevilicoGPT/Revilico Agent** as well for targeted protein sets)
**Workflow**
1. **Curate Disease-Relevant Protein Library**
Assemble a comprehensive library of protein structures associated with your therapeutic context to define the search space for reverse docking. Use **RevilicoGPT/Revilico Agent** to query structural databases (PDb, AlphaFold Database) and filter for disease relevant targets
Sample Query: I have a drug that shows efficacy in \[disease/phenotype]. Please compile a library of \~1000 protein structures associated with this disease, prioritizing druggable targets like kinases, GPCRs, enzymes and ion channels. Return PDB files, amino acid sequences, or AlphaFold structures with binding site annotations.
**Revilico Agent** will output a curated protein structure library with metadata including protein family, known functions, disease associations, and druggability scores. Users can refine this library by filtering for specific protein families or expanding to the full druggable proteome for unbiased screening. The library should include experimentally resolved structures when available and high-confidence AlphaFold predictions to maximize coverage. For proteins with multiple conformations available, include representative structures capturing different functional states to account for conformational selectivity
2. **Progressive Virtual Screening via Multi-Tiered Docking**
Screen your drug molecules against the protein library using a cascade of increasingly rigorous docking methods to efficiently identify and refine target candidates. Start with rapid screening to eliminate non-binders, then apply more accurate methods to top candidates.
Run **Static Docking** against all \~1000 proteins in your library. This high-throughput screen rapidly eliminates non-binders based on binding affinity predictions. This can be done through ‘blind docking methods’ where the entire protein surface is exposed to the algorithm to calculate binding energies, placing the compound in what it believes to be the best suited position. Downselect to top 50-100 proteins with binding affinities \< -8 kcal/mol, prioritizing targets with multiple low-energy poses in consistent binding sites, which indicates reproducible binding modes rather than spurious predictions
Run **Flexible Docking** on the top 50-100 candidates from static docking. This will account for side-chain flexibility in the binding site, providing more realistic binding geometries and hybrid physics ML scoring. Downselect to the top 20-30 proteins based on CNN Pose Score, CNN Affinity, and low intramolecular (sterics) energy.
For targets where protein flexibility is critical, you can optionally run **Ensemble Docking** using MD-derived conformational ensembles. This captures how your drug binds across the protein’s conformational landscape. Prioritize targets showing consistent binding across multiple protein conformations over those with conformation-dependent binding.
3. **Validate Binding Modes with Pharmacophore Analysis**
For the top 20-30 candidates from flexible docking, analyze the binding interactions to confirm chemically reasonable binding modes before committing to expensive MD simulations. Run **Pharmacophore Analysis** on the top-ranked docked poses. Prioritize targets where the drug forms multiple complementary interactions (typically 3-5 key contacts). Filter out targets where binding relies solely on hydrophobic burial with no directional interaction. Use **Conformer Search** on your drug to validate that the bound conformation is energetically accessible to ensure the drug doesn’t pay a large conformational penalty for binding.
4. **Dynamic Stability Assessment via Molecular Dynamics**
Run **Protein-Ligand MD** simulations for your top 3-5 target candidates. MD reveals whether static docking purposes remain stable or dissociate, whether key interactions persist, and whether the protein binding site accommodates the drug without creating a strain. Strong candidates show stable complex formation, maintained binding interactions, and no signs of ligand egress or binding site distortion.
5. **Quantitative Affinity Ranking via Free Energy Calculations**
For your top 2-3 candidates, calculate **absolute binding free energies** to quantitatively rank targets and identify the primary binding partner versus secondary off-targets. The target with the most favorable (most negative) ΔG binding is your primary molecular target, while targets with weaker binding may represent secondary pharmacology or off targets.. Consider running **MMPBSA** analysis from MD trajectories as a faster less rigorous alternative if FEP is computationally expensive and costs too many credits for your campaign.
**Results**
* Ranked candidate target list with binding affinities and pose clusters
* Pharmacophore interaction maps (e.g. H-bonds, salt bridges, key contacts)
* MD stability metrics (RMSD, RMSF, interaction persistence)
* Absolute binding free energies (ΔG bind) per target
* Primary target identification with validated binding mode
Look for convergent evidence across methods, your primary target should rank consistently high in docking scores, show 3-5 chemically reasonable interactions in pharmacophore analysis, maintain stable binding in MD simulations, and demonstrate the most favorable ΔG bind in FEP calculations. Use the validated primary target structure and binding mode to proceed with structure based lead optimization, while documenting secondary targets with moderate affinity for selectivity optimization or polypharmacology consideration. You should also keep in mind that your compound could be engaging with and interacting with several targets all at once, so you need to overlay your results with biologically feasible pathways that you believe are being affected during phenotypic screening. This helps to overlay and connect the dots between binding engagements and biological outcomes.
**Now what?** You may have your target of interest now (after you have validated it experimentally) and may want to run optimization to get tighter biochemical engagements
* After analyzing and down-selecting your targets of interest, you can then move into generative chemistry campaigns using De Novo Library Generation, Molecular Optimization, or Custom Model training to get new libraries optimized for engagement to your target.
* After getting results from generative chemistry, you can re-score the new library using different tools to get and expand your lead series to be synthesized and tested.
**Why Revilico?**
This workflow addresses the challenges of identifying the primary molecular target for a drug with demonstrated efficacy but an unknown mechanism of action, which is critical for rational structure based optimizations. This process systematically identifies and validates protein targets of interest before moving forward with costly experimental validations like siRNA knockdowns or CRISPR screens.
# Druggable Binding Site ID
Source: https://docs.revilico.bio/solutions/target-identification/Identifying-Druggable-Binding-Sites
Identifying Druggable Binding Sites on Challenging “Undruggable” Targets
**The Problem You are Trying to Solve:**\
*“I have a molecular target without obvious or stable binding sites, and I want to identify druggable pockets suitable for structure based drug design”*
Traditionally identifying binding sites on “undruggable” targets required expensive experimental methods like fragment screening via X-ray crystallography, NMR spectroscopy to detect conformational dynamics, or hydrogen-deuterium exchange mass spectrometry to identify flexible regions, all requiring months of work and significant material costs with no guarantee of finding a druggable site.
**Solution**\
This workflow enables users to a) systematically identify all potential binding sites on challenging targets through static pocket analysis and b) discover cryptic, transient, and allosteric pockets that only appear during protein motion using molecular dynamics-based pocket detection across Revilio’s structural analysis engines. Proper druggability assessment is achieved by analyzing pocket properties across multiple protein states and ranking sites by persistence, accessibility, and druggability scores, helping eliminate status analysis blind spots and increasing confidence that identified pockets represent genuine druggable sites even on targets traditionally considered undruggable.
**What Data Do I Need to Provide?**
* PDB File of your protein structure (will be used for Pocket Search, Docking, and MD simulations)
**Workflow**
1. **Baseline Static Pocket Identification**
Run **Pocket Search Engine** on your protein structure. If you have multiple conformational states, run the analysis on each independently after extracting trajectory .pdb files. This engine will identify cavities and rank them by druggability score, volume, and physicochemical properties. For traditionally undruggable targets, you may find a few pockets or only low scoring ones at this stage. Document all identified pockets with their location as these provide context for interpreting cryptic pockets discovered later.
2. **Sample Conformational Landscape via Molecular Dynamics**
Run **Protein-Water MD** simulations with explicit solvent for 100-200 ns. The simulation allows the protein to explore thermally accessible conformational states under physiological conditions. During the trajectory, pockets will open, close, expand, contract, and new cavities will transparently appear as the protein breathes and fluctuates. The MD trajectory becomes the input for temporal pocket analysis in Step 3. For particularly challenging targets, consider running multiple independent MD simulations starting from different initial conformations to enhance conformational sampling and increase probability of discovering rare cryptic pocket opening events. You can also explore new configurations of the solvent as well as longer time scales to allow the proteins to reach equilibrium.
3. **Identify Cryptic and Transient Pockets**
Run **MDPocket** on the MD trajectory from Step 2. MD pocket performs pocket detection at each trajectory frame and tracks pocket properties over time identifying cryptic pockets, transient pockets, persistent pockets, and pocket frequency. MDPocket will output ranked cryptic pockets with temporal profiles showing when each pocket opens, how long it remains accessible, and its properties over time. Cryptic pockets with high opening frequency and favorable druggability represent prime targets for allosteric or induced-fit drug design strategies that capitalize on protein flexibility.
4. **Multi Conformational Validation**
Extract representative protein conformations from the MD trajectory and validate pocket characteristics in specific functional states to confirm MDPocket findings and understand pocket architecture. From the MD trajectory, extract 5-10 representative conformations showing different pocket states: cryptic pockets in fully open, partially open, and closed states. For more extensive analysis, you can cluster trajectory frames by backbone RMSD/RMSF or pocket volume to identify structurally distinct conformations. Run **Pocket Search Engine** on each extracted conformation independently. Compare pocket properties across conformations to understand how pocket characteristics change with protein state. Identify the “optimal binding conformation” where the cryptic pocket exhibits maximum druggability. This becomes the reference structure for subsequent structure based drug design and optimization studies.
5. **Validate Druggability with Fragment Screening**
Run **Static Docking** and **Flexible Docking** with a diverse fragment library (around 100-500 small fragments), into the top-ranked cryptic pockets identified in steps 3-4. Use the optimal binding conformation extracted in step 4 as the receptor structure, or perform ensemble docking across multiple conformations if the pocket shows conformational flexibility. For the top fragment hits in cryptic pockets, run brief **Protein-Ligand MD** simulations to confirm the pocket remains stable when occupied and the fragment doesn’t induce pocket collapse or protein unfolding. Successful fragment binding validates that the cryptic pocket is a genuine druggable site suitable for hit-to-lead optimization. We will soon be coming out with an engine that allows for pocket search using fragment ‘soup’ that will allow us to run longer term molecular dynamic trajectories in the presence of fragment perturbations to elucidate unique pockets of interest.
**Results**
* Comprehensive pocket inventory with druggability scores and volumes
* Temporal pocket profiles showing opening frequency and persistence across MD simulations
* Representative protein conformations with pocket open/closed states
* Fragment screening validation results
* Ranked druggability site recommendations with strategic classifications
Prioritize cryptic pockets with convergent evidence: high opening frequency, favorable druggability scores, successful fragment binding validation, and maintained stability in ligand-bound MD simulations. These represent genuine druggable sites on previously undruggable targets. If dynamic analysis reveals persistent cryptic pockets that were invisible in static structures and validate with fragment screening, you’ve successfully identified tractable binding sites for structure-based drug design
**Now what?** After identifying your pockets of interest, you can then move into identifying hit molecules for your campaign.
* Utilize the Virtual Screening engine or De Novo Library Generation to begin your campaign and identify initial hits. Utilizing the knowledge gained from pocket identification, you can target your docking calculations on the box of choice and across conformational states to ensure you’re representing the biological system properly.
**Why Revilico?**\
This workflow addresses the challenge of identifying druggable sites on targets without obvious or stable binding pockets by combining static analysis with dynamic molecular dynamics (MD) simulations. It systematically discovers cryptic and transient binding sites through MD and MDPocket, then validates the most persistent and druggable pockets using fragment screening to establish genuine targets for downstream structure-based drug design.
# Cell Line Identification
Source: https://docs.revilico.bio/solutions/target-identification/Identifying-the-Best-Therapeutic-Cell-Line
Identifying the Best Therapeutic Cell Line for a Drug-Target Complex
**The Problem You are Trying to Solve:**\
*“I have a known drug (ligand) and a target (protein) complex with demonstrated efficacy, and I want to identify which cell line will show the best therapeutic response based on its molecular profile.”*
Traditionally, identifying optimal cell lines for drug-target complexes required expensive, time-consuming wet-lab screens across dozens of cell lines, with each requiring weeks of cell culture, drug treatment, and phenotypic assays to measure therapeutic response. This empirical approach provided limited mechanistic insight into why certain cell lines responded better, making it difficult to predict responsiveness in new contexts or optimize for specific biological outcomes. The lack of comprehensive transcriptome profiling meant researchers often missed critical regulatory networks and temporal dynamics that govern drug response, leading to suboptimal cell line selection and failed validation experiments.
**Solution**\
This workflow enables users to a) systematically identify cell lines with optimal therapeutic response potential through target expression profiling and b) comprehensively characterize the molecular mechanisms driving drug efficacy using multi-modal transcriptome analysis across Revilico’s scRNA-seq Analysis Engine. We leverage three complementary analytic workflows, DEG Analysis for response magnitude assessment, Automated Target ID for regulatory network mapping, and Temporal Omics for response kinetics to ensure complete characterization of cellular response dynamics.
**What Data Do I Need to Provide?**
* Protein name or identifier (Required to identify cell line candidates to evaluate)
**Workflow**
1. **Identify Cell Lines based on Target Expression**
Determine which cell lines express your target protein at therapeutically relevant levels. To do this on Revilico, users can leverage **Revilico Agent** to query expression databases and identify candidates based on normalized transcript abundance (nTPM scores)
Sample Query for General Screening: I have the protein AXL. I want to identify the top 20 cell lines based on nTPM scores.
Sample Query for Disease-Specific Screening: I have the protein AXL. I want to identify the top 20 cell lines associated with Triple Negative Breast Cancer based on nTPM scores.
**Revilico Agent** will output a ranked list of cell lines with corresponding expression levels. Users can cross-validate these results using the Human Protein Atlas Database by searching for the target protein and reviewing associated cell lines. High expression generally indicates the target is biologically active in that cellular context, making it a strong candidate for therapeutic response studies
2. **Acquire Paired scRNA-seq Datasets**
Obtain control and experimental scRNA-seq datasets for each candidate cell line that include your target gene. Users leverage Revilico Agent to source publicly available datasets in h5ad format, ensuring both baseline and treatment conditions are represented.
Sample Query: I have these cell lines that I have identified: \[list]. Please provide me a panel of scRNA-seq datasets in the format of h5ad files for each cell line. For each cell line, the h5ad files should come in pairs, one control sample and one experimental sample. Ensure that these h5ad files include our target gene as well.
Revilico Agent will output matched control-experimental h5ad file pairs for each cell line, verified to contain your target gene in the expression matrix. These paired datasets represent the transcriptomic state before and after perturbation, enabling comparative analysis of cellular response. If suitable public datasets are unavailable, user can query from Gene Expression Omnibus or have to generate custom scRNA-seq data, or refine their cell line selection criteria based on data availability.
3. **Run Comprehensive scRNA-seq Analysis**
Execute all three analytic workflows within Revilico’s **scRNA-seq Analysis engine** sequentially for each control-experimental h5ad pair using your selected or generated gene panel. This mullti-modal analysis provides complementary perspectives on therapeutic response.
The **DEG Analysis** will provide a response magnitude assessment. Here we will identify whether or how strongly each cell line responds at the transcriptomic level. Here volcano plots reveal statistical significant gene expression changes, clustering visualizations show distinct transcriptomic states between control and experimental conditions, and pathway enrichment analyses identify affected biological processes. To further understand how to interpret the outputted graphs, you can refer to the engine specific documentation here: scRNA seq Analysis engine
The **Automated Target ID** will provide a mechanistic pathway analysis. Here the engine constructs gene regulatory networks to reveal how the drug-target interaction affects cellular behavior through regulatory cascades. In additional phenotype predictions will be provided, and should align with your therapeutic goals. To further understand how to interpret the outputted graphs, you can refer to the engine specific documentation here: scRNA seq Analysis engine
The **Temporal Omics Engine** provides response kinetics modeling. Here the engine will analyze the temporal dynamics to understand how fast and in what sequence cellular response occurs. Pseudotime heatmaps order gene expression changes by progression, and ODE based simulations predict temporal dynamics of key genes.To further understand how to interpret the outputted graphs, you can refer to the engine specific documentation here: scRNA seq Analysis engine
**Results**
* List of candidate cell lines and their associated nTPM scores
* A panel of scRNA seq datasets corresponding to the candidate cell lines
* DEG Analysis plots (Volcano Plots, Clustering Visualizations, Differential expression gene lists)
* Automated Target ID plots (GRN Graphs, Cascade trees, phenotype bar graphs)
* Temporal Omics (Pseudotime heatmaps, Trajectory plots, ODE simulation results)
When interpreting, you will keep these key metrics in mind. These 4 metrics will hold the highest weight: nTPM scores, DEG count and fold change magnitude, cascade tree depth and breadth, and phenotype score alignment. High nTPM scores indicate more drug-target engagement, more significant DEGS indicate a stronger response, deeper and broader cascades indicate extensive effects, and phenotype predictions match therapeutic goals. With the cascade trees you can evaluate whether key genes are being affected downstream of your target gene. In addition you can use **Revilco Agent** to assist you in evaluating and ordering the cell lines based on these metrics
* **Why Revilico?**\
The goal of this workflow is to systematically identify the cell line that will show the best therapeutic response for a drug-target complex based on its molecular profile before moving into costly experimental tests. This is achieved by first identifying candidates based on target expression (nTPM scores) and then running a comprehensive multi-modal scRNA-seq analysis (DGE, Automated Target ID, Temporal Omics) to characterize the molecular mechanisms driving drug efficacy.
# Protein Variant / Mutagenesis Validation at Scale
Source: https://docs.revilico.bio/solutions/target-identification/Protein-Variant-Mutagenesis-Validation-at-Scale
## **The Problem You Are Trying to Solve**
*“I want to run large-scale separation-of-function mutants on my target to validate ligand–protein interactions, but experimentally generating and testing all variants is too intensive.”*
Mutagenesis is one of the most powerful tools for:
* Validating binding site hypotheses
* Confirming key interaction residues
* Distinguishing orthosteric vs allosteric mechanisms
* Understanding resistance mechanisms
However, experimentally generating these proteins and screening dozens or hundreds of amino acid substitutions (across multiple residues and ligands) quickly becomes infeasible.
**Solutions**\
This workflow simulates mutational effects on binding affinity, pose stability, and protein conformational behavior before any wet-lab investment. The primary workflow chain is as follows: Baseline Complex Validation → In Silico Mutant Generation → Docking Re-evaluation → MD Stability Assessment → ΔΔG / Free Energy Comparison → Ranked Variant Prioritization.
This enables users to computationally predict the effects of amino acid changes on small-molecule binding, allowing you to prioritize which variants to generate experimentally.
## **What Data Do I Need to Provide?**
Required
* Protein structure (PDB or predicted structure)
* Ligand structure (SMILES or complex PDB)
Recommended
* Known binding pose (co-crystal or validated docked model)
* List of residues of interest that contribute to biological activities or ligand engagement (e.g., predicted binding site residues)
Optional
* Known resistance mutations which can be based on pure biophysical analysis of the protein or based on genomic patient data to test for certain amino acid substitutions that exist in the general patient populations.
* Functional assay results for calibration
## **Workflow**
1. **Validate the Wild-Type Baseline**
Before evaluating mutations, ensure the starting complex is well-characterized. Users typically:
* Generate the baseline structure using the amino acid only across AlphaFold, OpenFold, and Boltz structure generation engines
* Confirm ligand binding pose with **Docking or Co-Folding models**
* Validate stability with short **Protein–Ligand MD**
* Identify key hydrogen bonds, electrostatic contacts, hydrophobic interactions and how they change as a function of time.
The goal of this step is to establish a high-confidence wild-type binding model to serve as the reference state. This provides baseline binding affinity metrics, stable binding trajectory, and identified “hotspot” residues.
2. **Generate In Silico Mutants**
Now introduce targeted amino acid substitutions. Users:
* Select residues for mutagenesis (e.g., binding site, distal regulatory regions)
* Define mutation types (alanine scan, conservative substitutions, resistance-inspired variants)
* Generate mutant protein structures computationally using our variety of protein structure generation tools
Mutations may include:
* Amino acids that are known drivers of binding and activity. Mutations can allow confirmation of key binding modes and energies.
* Alanine scanning (loss-of-function mapping)
* Charge swaps (e.g., Asp → Lys)
* Size/steric changes (e.g., Phe → Ala)
* Known clinical resistance mutations
This gives a panel of computationally generated mutant protein structures.
3. **Re-evaluate Ligand Binding via Docking**
For each mutant:
* Re-dock the ligand into the altered structure with **Static Docking or Co-Folding** as a first-pass
* Compare predicted binding affinity and pose shifts, and you can confirm certain binding poses and amino acid engagements using the **Pharmacophore Analysis Engine.**
* Repeat with **Flexible Docking** if local rearrangement is expected, or with Ensemble docking for more sophisticated and dynamic protein systems.
**Key evaluation:**
* Loss or gain of critical contacts
* Pose displacement
* Score degradation relative to wild-type protein mutants
This gives a first-pass ranking of mutations by predicted binding impact.
4. **Assess Dynamic Stability (Optional)**
Docking captures static geometry. Now evaluate dynamic effects. Users run:
* **Protein–Ligand MD** for high-impact mutants
* Compare RMSD, RMSF, SASA, and interaction persistence against wild-type
**Key questions:**
* Does the ligand remain stably bound?
* Does the mutation induce destabilizing pocket rearrangement?
* Does distal mutation alter global conformational behavior?
This provides stability-confirmed mutation impact assessment even across dynamic time conditions with differing solvent parameterizations.
5. **Quantify Binding Impact**
For prioritized mutants, estimate energetic consequences. Users can:
* Use **MMPBSA or MMGBSA on Protein Ligand MD** for rapid ΔG comparison
* Use **RBFE** to compute ΔΔG between wild-type and mutant complexes
* Use **ABFE** for absolute binding comparison (if needed)
What you measure:
* ΔΔG = ΔG\_mutant − ΔG\_wild-type
* Negative ΔΔG → improved binding
* Positive ΔΔG → weakened binding
This gives a quantitative ranking of mutational effects across different protein variants.
## **Results**
* Ranked mutation list by predicted impact on binding
* Identified critical binding residues
* Separation-of-function candidate mutants
* Energetic and structural rationale for each mutation
This enables:
* Focused experimental validation
* Reduced mutagenesis burden
* Mechanistic clarity for ligand engagement
## **Integration with Other Engines (Optional)**
Depending on study goals, this workflow can integrate with:
* **Allosteric Design workflows** (if distal mutations alter regulatory pockets)
* **QSAR Modeling** (incorporate mutation-dependent activity shifts)
* **ADMET-AI** (if mutations influence ligand orientation and downstream design)
* **scRNA-Seq Analysis** (if variant alters pathway activation)
* **Quantum Chemistry engines** (if mutation changes electrostatic environment significantly)
**Now what?**
* You can utilize this information to create more informed protein variants of your validated complexes that you’d like to evaluate for separation of function mutagenesis.
* Creating the right complexes can allow you to confirm structural hypotheses of binding, and to ensure that your hypotheses of what amino acids drive binding are correct
* You can utilize a variety of different engines on the dashboard to test your new protein variants for structural similarity, or your new complexes in more advanced scoring engines.
## **Why Revilico?**
Revilico enables scalable mutational validation by combining:
* Structural modeling of variants
* Docking-based interaction reassessment
* Dynamic stability simulation
* Thermodynamic free energy comparison
* Unified visualization and interpretation
Instead of experimentally generating hundreds of mutants, you computationally triage variants and focus wet-lab efforts on those most likely to reveal meaningful mechanistic insights.
# Disease Bio and Mechanisms
Source: https://docs.revilico.bio/solutions/target-identification/Understanding Disease-Biology-and-Mechanisms-on-Revilico-OS
Understanding Disease Biology and Mechanisms on Revilico OS
**The Problem You are Trying to Solve:**\
*“I want to understand how my disease of interest works biologically (what pathways, mechanisms, and cellular programs are driving it) before committing to target selection or downstream drug discovery workflows.”*
Before pursuing structure prediction, molecular docking, or *de novo* compound design, it is critical to understand the biological context of a disease. This includes identifying dysregulated pathways, cell-type-specific transcriptional programs, and temporal dynamics that underlie disease onset and progression. However, extracting these insights from high-dimensional omics data is often nontrivial and requires specialized bioinformatics expertise, custom pipelines, and significant interpretation effort. A centralized, interpretable workflow that connects differential expression, target prioritization, and temporal biology enables researchers to move from raw data to mechanistic understanding with confidence.
**Solution**\
Revilico provides a unified biological discovery layer designed to help users understand disease mechanisms and pathways at the cellular level. Using transcriptomics-driven target identification workflows, the platform allows users to interrogate how a disease operates biologically before advancing to structural biology or chemistry-driven pipelines. This disease-understanding layer can serve as a preceding step to downstream target validation and structure-based discovery, enabling users to ensure that their hypotheses are grounded in biological signal rather than isolated assumptions. In addition to data-driven analyses, users may optionally leverage RevilicoGPT as a contextual intelligence layer to synthesize known biology, pathway annotations, and mechanistic hypotheses alongside their experimental results.
**What Data Do I Need to Provide?**
* Control and Experimental h5ad files
* Bring some understanding of your target disease of interest
**Workflow**
1. **Differential Gene Expression (DGE) → Identify What is Changing**
The first step in understanding a disease is determining what molecular programs are altered relative to a control or baseline state. Users upload control and experimental scRNA-seq datasets (h5ad files). Revilico computes differential gene expression across relevant cell populations, identifying genes that are significantly upregulated or downregulated in the disease or experimental condition.
To interpret the results:
* Upregulated genes often indicate activated pathways, stress responses, or compensatory mechanisms
* Downregulated genes may reflect loss of normal function, differentiation defects, or pathway suppression
* Patterns across specific cell types help distinguish cell-intrinsic drivers from systemic effects
At this stage, the goal is not to select a final target, but to form an initial hypothesis about which biological processes are disrupted.
2. **Automated Target Identification → Prioritize What Matters**
Once transcriptional changes are identified, the next challenge is determining which genes are most biologically and therapeutically relevant. Revilico’s Automated Target ID layer ranks candidate genes emerging from DGE by integrating:
* Magnitude and consistency of expression change
* Cell-type specificity and relevance
* Signal robustness across samples or conditions
This step reduces hundreds or thousands of DE genes into a shortlist of biologically meaningful targets.
To interpret the results:
* High-ranking targets are likely to be drivers rather than passengers
* Cell-type-restricted targets may suggest precision intervention opportunities
* Broadly dysregulated targets may indicate core disease machinery
Rather than forcing a single “best” answer, this step provides a ranked landscape of plausible intervention points.
3. **Temporal Omics Analysis → Understand How the Disease Evolves**
Many diseases are dynamic systems. Understanding when biological changes occur is as important as knowing what changes. For datasets with temporal structure (e.g., disease progression, treatment timepoints), Revilico models gene expression as a function of time. This reveals dynamic patterns in pathway activation and target behavior.
To interpret the results:
* Early-changing genes may represent initiators or causal drivers
* Late-changing genes may reflect downstream effects or compensatory responses
* Transient expression peaks can indicate regulatory switches or checkpoints
Temporal insights help differentiate targets that are merely associated with disease from those that may be strategically actionable depending on intervention timing.
**Results**
* Key dysregulated genes and pathways associated with the disease
* Prioritized target candidates grounded in transcriptomic evidence
* Cell-type-specific and temporal insights into disease mechanisms
* A biologically informed hypothesis of how the disease operates at the molecular level.
These results collectively provide a mechanistic understanding of the disease, rather than a single-point target guess. RevilicoGPT can be used alongside these analyses to contextualize identified genes with known pathways and disease biology, summarize mechanistic hypotheses supported by both literature and data, and assist in interpreting complex multi-gene or temporal patterns. This added layer complements experimental data without replacing it, helping users reason about why observed patterns may be occurring.
**Now what?** Now that you have a strong understanding of your disease biology and candidate targets:
* If you want to validate or explore the target structure → proceed to AlphaFold or OpenFold within our platform.
* If you have compounds or libraries → advance to binding chemistry or quantum chemistry workflows.
* If you do not yet have compounds → use *de novo* library generation to design molecules for your selected target, or make use of our on hand liquid stock or powder libraries to be delivered after structure based drug design has been concluded.
**Why Revilico?**\
This Revilico workflow enables users to query the availability of experimental data before proceeding to AI structure prediction. Utilizing these transcriptomic based workflows allows you to understand the deeper mechanics driving your disease of interest to investigate or validate therapeutic avenues and proteins to target.
# 2D/3D Structure MedChem Review with Poses
Source: https://docs.revilico.bio/tutorials/medchem-review/2d-3d-structure-medchem-review-with-poses
Downselect compounds from a high throughput screen and triage them — poses and all — into a collaborative RevStudio Compound Review session
## Overview
Once you've run a high throughput virtual screen, the next step is to analyze the results at scale and downselect your compounds. This tutorial covers that full downselection workflow in **RevStudio's Compound Review** — starting from your **RevScreen** (static, flexible, or ensemble docking) or **Boltz2** results, and going all the way through to a 3D pose-level review that your medicinal chemistry team can use to sign off before compounds move to the wet lab.
Unlike a purely 2D triage, this workflow pulls the actual **top-ranked 3D poses** for each compound directly from the pipeline they were generated in, so your team can review binding mode alongside 2D structure in a single collaborative session.
[Downselecting Compounds After High Throughput Screening — Watch Video](https://www.loom.com/share/abc3dd764f8e483f91f6b90b2a37ff29)
***
## When to Use This Workflow
Use this workflow any time you've run one or more screening campaigns or batches and need to downselect to a smaller, high-confidence compound set — with 3D poses in hand — before moving to review or synthesis:
* **RevScreen — Static & Flexible Docking**
* **RevScreen — Ensemble Docking**
* **Boltz2 Cofolding**
If you've run multiple campaigns or batches, you can select and download several at once, run your own analysis externally (or in RevAgent), and then bring your downselected list back into RevStudio for 3D-aware review.
***
## Step 1: Export and Downselect Your Results
From the Engines tab, open **RevScreen** (or **Boltz2**) and go to the analysis view for your campaign.
If you've run multiple campaigns or batches, select the ones you want to pull results from. You can download and combine several batches at once.
Export the compound activity metrics, confidence scores, and any other data you need from RevScreen or Boltz2 as a CSV.
Concatenate your exported CSVs if needed and run your own filtering — the goal is to land on a downselected list, typically somewhere in the range of 50–200 compounds, that's ready for downstream review.
The only field you strictly need in your final CSV is a column of SMILES strings. Login IDs, affinities, pose counts, and other metrics can all ride along in the same file.
***
## Step 2: Create a Compound Review Session with Poses
With your downselected CSV ready, head to **RevStudio → Compound Review** to bring the compounds — and their 3D poses — into a shared review session.
If you only need to review compounds at the 2D level, you can upload your CSV directly. If you want to pull the actual 3D poses generated for each compound, use **Import from Pipeline** instead.
Click **Import from Pipeline** and create a new review. Give it a clear name and, optionally, a description.
You can filter based on top percentage of performers overall, top percentage within the pipeline, or a fixed top number of compounds — downselection thresholds are calculated on the docking results (e.g. static/flexible ensemble docking affinities).
Choose the campaign or pipeline your compounds came from.
Upload the SMILES CSV with your associated data, then click **Search Pipeline**. RevStudio will match your compounds against the full pipeline and report how many matched (for example, 487 of 800 compounds).
Click **Create Review**. RevStudio pulls all matched compounds — along with their top-ranked poses — into a structured review session.
***
## Step 3: Review Compounds in 2D and 3D
Inside a review session, you can switch between a **tabular view** and a **grid view**. Grid view is the most common choice for stepping through compounds one at a time with full context visible.
### Working Through the Compound Set
Use the grid view to click through each compound and review its data, 2D structure, and associated metrics.
Click **Share** to give teammates in your organization access. Shared sessions appear under their **Shared with Me** tab, and all changes either of you make are reflected in the same collective session.
Comment directly on a compound to record your assessment. Comments are attributed to their author, and you can edit or delete your own comments as the review progresses.
Flag compounds you and your collaborators want to downselect further, then filter the session to show only flagged compounds.
Filter on pose count or any other computed metric to narrow down the set you're actively evaluating. Filtered results update both the grid and the 2D structure side panel.
### Reviewing the 3D Pose
Each compound imported via pipeline comes with an associated 3D profile — the pose pulled from the source campaign.
Click **View 3D Pose** to open a larger, more refined view of the binding pose, including core interactions — whether the pose came from Boltz2 or a static/flexible ensemble docking pipeline.
Pop the 3D viewer out to expand it further and get a clearer look at the interactions.
Currently, only one pose is shown per compound. Support for viewing all of an engine's top nine generated poses is planned.
While reviewing the 3D pose, you can still write notes, like/dislike the compound, and add tags — for example, flagging that a compound engages certain amino acids, or that it's non-synthesizable.
***
## Step 4: Configure Tags
Tags help you categorize compounds against structural or mechanistic criteria as you review.
Use the tag panel to create, edit, or turn off tags.
Assign colors and icons to tags so they're easy to scan at a glance across the grid view.
***
## Step 5: Head-to-Head Compound Comparison
When you need to compare a handful of compounds side by side:
Choose the compounds you want to compare — for example, three or six candidates you're deciding between.
Pull them all into the **Compare** tab to view them simultaneously and make your assessment across the set.
Click **Back to Review** at any point to return to the full compound list and continue your session.
***
## Why This Matters
This workflow turns a large, multi-batch screening output into a tight, well-documented compound set ready for the wet lab:
* Downselect from hundreds of thousands of screened compounds down to a manageable shortlist using your own filtering criteria
* Bring 2D structure and top-ranked 3D poses into the same collaborative session, so reviewers aren't switching tools to judge binding mode
* Capture flags, tags, comments, and likes/dislikes as a structured, shared record of *why* a compound advanced or didn't
* Use head-to-head comparison to make final calls across your top candidates before triaging to the wet lab in **RevLab**
***
## Next Steps
* [**2D Structure MedChem Review**](/tutorials/medchem-review/2d-structure-medchem-review) — The 2D-only version of this workflow, for when 3D poses aren't needed
* [**Compound Review Reference**](/docs/revcompound-review) — Full reference documentation for RevStudio's Compound Review engine
* [**RevScreen - Static & Flexible Docking**](/tutorials/rev-bind/static-flexible-docking) — Generate the docking results used in this workflow
* [**RevScreen - Ensemble Docking**](/tutorials/rev-bind/ensemble-docking) — Generate ensemble docking results for pipeline import
* [**Boltz2 Cofolding**](/tutorials/rev-bind/boltz2-cofolding) — Generate co-folded poses for pipeline import
# 2D Structure MedChem Review
Source: https://docs.revilico.bio/tutorials/medchem-review/2d-structure-medchem-review
Export RevBind screening results and triage them into a collaborative RevStudio Compound Review session for medicinal chemistry sign-off
## Overview
Once a **RevBind** screen has finished running — whether through **RevDock** (Static, Flexible, or Ensemble Docking) or **Boltz2 Cofolding** — you'll have a set of poses and computed activity metrics for every compound in your library. Before those compounds move forward to synthesis and testing, they typically need a second set of eyes: a medicinal chemist who can review the 2D structures, flag problematic functional groups, and sign off on what actually gets made.
This tutorial walks through exporting your RevBind results and triaging them into a **RevStudio Compound Review** session, where your MedChem team can review, annotate, tag, and approve compounds in a shared, structured workspace.
[Export RevBind Results for MedChem Review — Watch Video](https://www.loom.com/share/2e24a04d8aa84cfba70cfec1a435d6a2)
***
## When to Use This Workflow
This handoff pattern applies any time you've run a RevBind screen and need medicinal chemistry input before moving forward, regardless of which engine produced the results:
* **RevDock — Static Docking**
* **RevDock — Flexible Docking**
* **RevDock — Ensemble Docking**
* **Boltz2 Cofolding**
In every case, the export and triage steps are identical — only the source of the compound-by-compound metrics changes.
***
## Step 1: Export Your RevBind Results
After your RevBind screen completes, take a look at the poses and analytics in the results viewer to get a general sense of the run before exporting.
Below the 3D viewer and summary analytics, scroll down to the compound-by-compound data table. This gives you a per-compound breakdown of every metric the engine calculated — docking scores, confidence, predicted activity, and so on.
Click **Export CSV** and download the file.
Open the CSV and rename the column containing your molecular structures to **SMILES**. Your file should end up with a `SMILES` column containing the SMILES strings for every compound, alongside all of the computed metrics from your run.
RevStudio auto-detects common header variants such as `smiles`, `SMILES`, and `smilesStrings`, so exact casing isn't critical — but keeping it clean and consistent makes the file easier to work with downstream.
***
## Step 2: Create a Compound Review Session
With your export in hand, the next step is to triage it to your medicinal chemistry team in **RevStudio**.
From the Revilico OS dashboard, go to **RevStudio → Compound Review**.
Give the session a clear, descriptive run name (e.g., `AXL-medchem-review`).
Use the notes field to capture the purpose of the review — for example, noting the target and the fact that the set needs review and approval before moving to synthesis (e.g., *"AXL target, needs review and approval before sending to synthesis"*).
Drag and drop the CSV file you exported and renamed in Step 1 into the upload area. RevStudio will detect the SMILES column automatically and generate 2D structure renderings for every compound in the set.
Once processed, you'll see your new review session listed with the compound count, session name, and creation date.
***
## Step 3: Share the Review Session
If your medicinal chemistry team needs access, share the session directly from RevStudio:
Click **Share** on your review session.
Share with the relevant user(s) and set their access as **Read** or **Write**. For internal team review, **Read** access is usually sufficient unless collaborators need to add their own annotations.
Once shared, close the dialog — the session is now available to your teammates under their **Shared with Me** tab.
***
## Step 4: Review Compounds
Inside a review session, you can switch between two layouts:
* **Card view** — the default and most commonly used layout for stepping through compounds one at a time
* **Tabular view** — a spreadsheet-style layout for scanning the full compound list at once
### Working Through the Compound Set
Once all compounds are loaded, list them out and click through each one to review its 2D structure and associated metrics.
Freeze the left or right-hand side panel if you need a consistent reference point while scrolling through structures and data.
Use the like/dislike controls to record a quick directional call on each compound.
Leave a comment on any compound to explain your reasoning — for example, noting that a functional group may not be feasible for the campaign, or that a substituent could introduce toxicity or metabolic instability.
Flag compounds that need special attention or follow-up before moving forward.
Select or create tags to categorize compounds against structural or mechanistic criteria — for example, tagging compounds expected to engage a specific residue (e.g., **Lys787**). Tags can be created, modified, and deleted from the tag management panel.
RevStudio's 3D structure visualizer will also surface within Compound Review sessions, letting you cross-reference the same tagging and flagging workflow against the 3D binding pose alongside the 2D structure.
***
## Accessing Shared Reviews
If a colleague has shared a review session with you, it will appear under the **Shared with Me** tab. Selecting a shared session gives you the exact same review tools described above — commenting, flagging, liking/disliking, and tagging — so your whole team works from one consistent set of annotations.
***
## Why This Matters
Compound Review turns a static CSV export into a living, structured checkpoint between computational screening and synthesis. Rather than passing spreadsheets back and forth over email, your medicinal chemistry team gets:
* A shared, versioned view of every compound coming out of a RevBind screen
* Structured annotations (flags, tags, comments) that capture *why* a decision was made, not just what the decision was
* A single source of truth your team can reference before compounds are queued for synthesis and testing
***
## Next Steps
* [**Compound Review Reference**](/docs/revcompound-review) — Full reference documentation for RevStudio's Compound Review engine
* [**RevScreen - Static & Flexible Docking**](/tutorials/rev-bind/static-flexible-docking) — Generate RevBind results using classical docking
* [**RevScreen - Ensemble Docking**](/tutorials/rev-bind/ensemble-docking) — Generate RevBind results using MD-sampled protein conformations
* [**Boltz2 Cofolding**](/tutorials/rev-bind/boltz2-cofolding) — Generate RevBind results using AI co-folding
# Managing Pipeline Projects Effectively
Source: https://docs.revilico.bio/tutorials/medchem-review/managing-pipeline-projects-effectively
Organize pipelines, notes, and files by target using Revilico OS Projects, and track campaign status on a shared Kanban board with timeline view
## Overview
As your research scales across multiple targets, keeping track of which pipelines belong to which campaign — and where each one stands — can get messy fast. **Projects** in Revilico OS give you a dedicated hub for organizing everything related to a single target or initiative: pipelines, notes, files, and campaign status, all in one place.
This tutorial walks through creating a project, organizing pipelines and notes within it, and using the built-in Kanban board to track your campaign status from kickoff to completion.
[Managing Research Projects and Campaign Status — Watch Video](https://www.loom.com/share/25da5043c86a4443921c163936f9419c)
***
## When to Use This Workflow
Use Projects any time you're running research against multiple targets or initiatives and need a single place to organize the pipelines, files, and notes tied to each one — for example, keeping an **EGFR** project separate from a **BRCA1** project, each with its own campaigns, documentation, and status tracking.
***
## Step 1: Create a Project
Create a project and give it a clear name — typically the target or initiative it represents (e.g., `EGFR`, `BRCA1`, or a working name like `Testing` for a demo).
Once created, this project becomes the hub for everything tied to that target — pipelines, campaigns, notes, and files all live here going forward.
***
## Step 2: Organize Pipelines Within a Project
From within a project, use the command center to drag and drop in the pipelines or campaigns you want to maintain under that project.
Each pipeline entry shows details like the pipeline type (e.g., static docking screen), its name, and when it was conducted.
Click into any pipeline to view it directly.
Click **Projects** to see and click through every project and its associated reviews. Selecting a project pops its data out into the central panel.
***
## Step 3: Share Projects and Manage Notes
Share a project with your team to give everyone an overview of all the files and pipelines associated with it.
Create new notes or folders within the project, or drag and drop existing notes in directly, to keep documentation and context alongside the work itself.
Drag and drop data associated with the target into the project through the data editor panel.
***
## Step 4: Track Campaign Status with the Project Board
Each project includes a Kanban-style board for tracking the status of every campaign tied to that target.
Click to edit the project. Here you'll see all of the campaigns tied to it, where you can download data, view analytics, or add more pipelines.
Add a card for a screen you're planning — for example, a static docking screen against your target of interest. Give it a name (e.g., `Static Docking Screening`) and set its status to **To Do**.
Use the card description to log the campaign's background, notes, and protocols — for example, noting that you're running an EGFR screening campaign with two million compounds from a specific library (e.g., Enamine's REAL library) — so the rest of your team has full context.
Select the engine the campaign will use (e.g., RevPocket) and assign the card to a team member.
Create the card. It can now be moved around the board by your team as work progresses.
Once a screen has been run, attach its pipeline to the card — the pipeline doesn't need to be fully processed; it can be attached at any stage. This keeps the card's status tied directly to the underlying pipeline.
You can also add a card for a pipeline that already exists within the project — for example, labeling it `Test Again`, setting a target date, and assigning it to yourself or a teammate. These display slightly differently from prospective cards, since they represent completed runs rather than planned ones.
Move cards across statuses (e.g., **To Do → In Progress → Done**) as work is completed, giving your whole team a live, shared record of what's been done and what's still outstanding.
Switch to the timeline view to see campaign dates laid out visually and stay on track with deadlines across the project.
***
## Why This Matters
Projects turn a scattered set of pipelines, files, and notes into a single, organized campaign hub:
* One place to find every pipeline, note, and file tied to a given target
* A shared Kanban board that keeps your whole team aligned on what's planned, in progress, and done
* A running log of campaign background, protocols, and decisions — not just raw pipeline outputs
* A timeline view for staying on schedule across multiple concurrent campaigns
***
## Next Steps
* [**2D Structure MedChem Review**](/tutorials/medchem-review/2d-structure-medchem-review) — Triage RevBind results into a collaborative Compound Review session
* [**2D/3D Structure MedChem Review with Poses**](/tutorials/medchem-review/2d-3d-structure-medchem-review-with-poses) — Downselect compounds with 3D pose-level review
* [**RevScreen - Static & Flexible Docking**](/tutorials/rev-bind/static-flexible-docking) — Generate the pipelines you'll track within a project
# Boltz2 Cofolding
Source: https://docs.revilico.bio/tutorials/rev-bind/boltz2-cofolding
Predict binding poses and activity metrics from a protein sequence and SMILES string using Boltz2's AI co-folding engine
## Overview
**Boltz2 Cofolding** is Rev-Bind's AI-powered structure prediction engine. Unlike classical docking, which requires a pre-determined 3D protein structure and a defined binding pocket, Boltz2 takes just two inputs — a **protein sequence** and a **SMILES string** — and generates a co-folded 3D structure showing the ligand placed within its most biologically relevant binding pocket.
This makes Boltz2 an especially powerful early-stage tool: you can generate binding pose predictions and activity metrics without needing a crystal structure or experimental pocket data.
[Running Boltz2 Protein-Ligand Cofolding Pipeline — Watch Video](https://www.loom.com/share/1b00be061f524e3ba446ba122ec24e99)
***
## How Boltz2 Cofolding Works
Boltz2 runs an AI engine that concatenates the protein sequence and small molecule structure, then simultaneously folds and docks them together. The resulting **co-folded structure** is a prediction of the binding pose — specifically, how the ligand sits within the pocket that drives the target's biological activity.
For example, with a kinase target the engine will co-fold the ligand into the ATP-competitive hinge region pocket — placing the compound where it would need to bind to produce a therapeutic effect — without requiring you to first define that pocket manually.
### Comparison to Classical Docking
| | Classical Docking (RevScreen) | Boltz2 Cofolding |
| -------------------------- | --------------------------------------------- | ------------------------------------------------ |
| Input | 3D protein structure + ligand 3D conformation | Protein sequence + SMILES |
| Pocket definition | Required (manual or RevPocket) | Automatic |
| Conformational flexibility | Limited | Full co-fold |
| Best for | Large-scale library screening | Early-stage pose prediction, novel targets |
| Output | Docking scores, poses | Co-folded structure, confidence, affinity, pIC50 |
***
## Output Metrics
Each co-folded result returns three primary metrics:
| Metric | Description |
| ------------------------ | -------------------------------------------------------------------------------------------- |
| **Confidence Score** | The model’s confidence in the predicted binding pose — reflects structural reliability |
| **Affinity Probability** | A heuristic estimate of binding affinity — useful for relative ranking across a compound set |
| **Predicted pIC50** | The primary activity outcome — a predicted potency metric derived from the co-folded pose |
Click the **i** icon next to any metric in the results panel for a plain-language explanation of what it means and how to interpret it.
These metrics together give you a rapid, hypothesis-generating view of how well a compound is likely to engage your target — useful for triaging compound sets before committing to more compute-intensive methods.
***
## Running a Boltz2 Pipeline
From the Revilico OS dashboard, go to **Rev-Bind → Boltz2**.
Click into **Data Engineering**. You’ll see two input fields: one for the **protein sequence** and one for the **SMILES string**.
If you don’t have a campaign running yet, use the **demo files** to follow along — they contain all the required inputs pre-formatted.
The protein sequence should be provided as a **CSV file** with a single chain column (e.g., `AXLChainA`) containing only the raw amino acid sequence string.
You have two options:
* **Drag-drop** your CSV file directly into the data engineering pane
* Click **Add Manually** and paste the sequence string directly into the input field
The demo file format shows exactly how the CSV should be structured. Reference it if you are building your own input file for the first time.
The small molecule is defined by its SMILES string. You can provide this as:
* A **CSV file** containing one or more SMILES strings for batch cofolding
* **Manual entry** — paste a single SMILES string directly into the input field
For a bulk campaign (e.g., 1,000 compounds), prepare a CSV with one SMILES per row and drag-drop it in.
Enter a descriptive name for the pipeline (e.g., `AXL-boltz2-type2-inhibitors`). This name will identify the run in your pipeline history.
Set the **Pipeline Type** to **Affinity** for standard protein-ligand cofolding. Additional co-folding types are available for more complex analyses:
* **Affinity** — Protein + ligand cofolding for binding pose and activity prediction *(standard)*
* **Multimer** — Co-fold multiple protein chains together
* Additional types for varied protein-small molecule configurations
Make sure to select the pipeline type that matches your experimental setup. Using the wrong type will produce mismatched output.
Click **Create Pipeline**. The engine will queue and run your co-folding job. You’ll see a confirmation message when the pipeline is created successfully.
Track all active and completed pipelines from the **central pipeline hub** in the navigation bar.
***
## Reading and Navigating Your Results
Once the pipeline completes, navigate to your results from the pipeline hub.
### Viewing Co-folded Structures
The primary output is an interactive 3D viewer showing the co-folded protein-ligand complex. The ligand will be placed within the highest-fidelity pocket — the site most likely to drive biological activity for that target class.
For kinase targets, for example, you’ll see the compound positioned within the **hinge region** — competing with ATP in the canonical ATP-binding site, which is exactly the mechanism a type II competitive inhibitor would engage.
### Navigating Large Result Sets
When running campaigns with many compounds, use the results panel to browse your full ligand list. You can jump directly to any specific compound by number (e.g., **Ligand 450**) rather than scrolling through the entire set.
### Exporting Data
For downstream analysis of large batches, click **Export CSV** to download your full results table. This includes all confidence scores, affinity probabilities, and predicted pIC50 values for every compound in the run — ready for statistical analysis, filtering, or integration into your own workflows.
***
## Calibrating Boltz2 Outputs
Boltz2's predicted metrics are powerful heuristics for **relative ranking** within a compound set. They are not absolute experimental measurements.
Best practice is to:
1. Run Boltz2 on a set of compounds that includes known actives and known inactives from your target
2. Validate that the predicted pIC50 and confidence scores rank-order consistently with your experimental data
3. Use this calibration to interpret predictions for novel compounds with greater confidence
Boltz2 outputs should always be calibrated against experimental results at each stage of your campaign. Treat predicted pIC50 as a relative ranking tool rather than an absolute potency prediction until calibrated for your target system.
***
## Triaging to Protein-Ligand MD
Once you have identified your highest-confidence co-folded poses, you can triage them directly into **Protein-Ligand MD** for more robust validation.
From the results panel, select your compound(s) of interest and click **Triage to Protein-Ligand MD**. This will:
* Export the co-folded pose as the starting structure for MD
* Queue an MD simulation to assess binding stability and residence time
* Enable MMPBSA and MMGBSA rescoring for more rigorous free energy estimates
This workflow is covered in depth in the [**RevMD-Bind tutorial**](/docs/revmd-bind).
***
## Next Steps
After Boltz2 cofolding, your compound list is ready for:
* [**Protein-Ligand MD — RevMD-Bind**](/docs/revmd-bind) — Triage top poses directly to MD for stability validation and MMPBSA/MMGBSA rescoring
* [**RevFEP**](/docs/revfep) — Rigorous alchemical free energy calculations for final-stage candidate ranking
* [**RevScreen - Static & Flexible Docking**](/tutorials/rev-bind/static-flexible-docking) — Run classical docking screens on the same compound set for orthogonal validation
* [**RevScreen - Ensemble Docking**](/tutorials/rev-bind/ensemble-docking) — Use MD-sampled protein conformations for maximum-accuracy virtual screening on shortlisted leads
# RevScreen - Ensemble Docking
Source: https://docs.revilico.bio/tutorials/rev-bind/ensemble-docking
A step-by-step walkthrough of Ensemble Docking in Rev-Bind — using molecular dynamics-sampled protein conformations for high-accuracy virtual screening
## Overview
Ensemble docking is the highest-accuracy docking mode in **Rev-Bind's Virtual Screening Engine**. Rather than docking against a single static protein structure, it samples the protein's dynamic conformational landscape via molecular dynamics (MD) simulation, then docks your compound library against multiple protein snapshots simultaneously. The result is a rich matrix of binding data that captures how compounds interact across the full dynamic behavior of the target.
[Ensemble Docking Guide — Watch Video](https://www.loom.com/share/96152485df7a4179b2a4e946f92294ea)
***
## Why Ensemble Docking?
| | Static / Flexible Docking | Ensemble Docking |
| ----------------------- | ------------------------------- | ---------------------------- |
| Protein representation | Single structure | Multiple MD snapshots |
| Conformational sampling | Limited to side-chain flex | Full protein dynamics |
| Accuracy | Good | High |
| Compute time | Fast | Longer |
| Best use case | Library-scale initial screening | Lead compound prioritization |
The trade-off is clear: ensemble docking takes more time, but it provides a substantially more accurate and dynamic picture of protein-ligand interactions. Accuracy typically translates far better than static docking, particularly for flexible or allosteric targets.
***
## Workflow Overview
Ensemble docking on Revilico OS involves three sequential stages:
1. **Run a Protein-Water MD Simulation** — Sample the protein's conformational space over time
2. **Pre-process the MD Trajectory** — Extract protein snapshots at defined time intervals
3. **Run Ensemble Docking** — Dock your compound library against each snapshot and aggregate results
***
## Stage 1: Run a Protein-Water MD Simulation
Before you can run ensemble docking, you need an MD trajectory to draw protein conformations from. This is done using Revilico's **RevMD-Aqua** engine.
From the Revilico OS dashboard, go to **Dynamic Molecular Interactions → Protein Water MD**.
Give your simulation a descriptive name (e.g., `AXL-apo-50ns`) to make it easy to identify when you return to ensemble docking.
Upload your protein structure PDB file, or drag-drop it from the file pane on the right-hand side of the screen.
Select your force field parameters. If you are new to molecular dynamics, the **recommended defaults** are appropriate for most standard protein systems.
For specialized systems, you can additionally configure:
* **Salt and ion concentration** — to match physiological or specific assay conditions
* **pH** — to reflect your experimental environment
Use the **Revilico Interpreter** to get plain-language explanations of any parameter on screen. This is particularly useful if you are newer to molecular dynamics setup.
Define your simulation time in nanoseconds. General guidance:
* **Minimum:** 50 ns for most protein targets
* **Recommended:** Scale upward based on protein size and the expected timescale of conformational transitions
* Longer simulations capture slower motions but require more compute time
Click **Run Simulation**. The platform will execute the protein-water MD job. You will receive a notification when the trajectory is ready. You can find a dedicated walkthrough video for RevMD-Aqua in the RevMD tutorials.
***
## Stage 2: Pre-process the MD Trajectory
Once your MD simulation is complete, return to the Virtual Screening Engine to prepare the protein snapshots for docking.
From the **Virtual Screening Engine**, select **Ensemble Docking**, then click **Pre-process MD**.
From the dropdown, select the MD simulation you completed in Stage 1. The platform will automatically preload the **min** (start time) and **max** (end time) values from your trajectory.
Adjust the time sliders to set which portion of the trajectory you want to sample from:
* **Full trajectory:** Leave the sliders at the preloaded min and max values to sample conformations across the entire simulation
* **Latter half only:** Advance the start slider forward to skip the early equilibration phase and focus on post-equilibrium conformations, which are generally more biologically representative
Sampling the latter half of the trajectory — after the protein has equilibrated — typically produces higher-quality ensemble members. Early frames may still reflect the starting crystal structure rather than native dynamics.
Define the time interval (in ns) between snapshots. The number of snapshots extracted is:
> **Number of snapshots = (max time − min time) ÷ interval**
**Example:** A 0–50 ns trajectory with an interval of 10 ns will produce 5 snapshots — one at the start and one at each 10 ns increment.
More snapshots increase the diversity and representativeness of the ensemble, but proportionally increase docking compute time. A starting range of **5–10 snapshots** is practical for most campaigns.
Click **Start Preprocessing**. The platform will extract protein PDB snapshots at each interval and prepare them for the docking stage. These pre-processed pipelines will appear in the ensemble docking interface.
***
## Stage 3: Run Ensemble Docking
With your pre-processed snapshots ready, you can now run the ensemble docking campaign against all protein conformations simultaneously.
In the Virtual Screening Engine, select **Run Ensemble Docking**.
From the dropdown menus:
1. Select the **MD simulation pipeline** you ran in Stage 1
2. Select the **pre-processed pipeline** you created in Stage 2 (the one containing your defined snapshot interval and sampling window)
You will be left with a set of PDB files — one for each protein conformation snapshot.
This is the critical step unique to ensemble docking. Because the protein moves and flexes across snapshots, the binding pocket position shifts with each conformation — so the docking box must be defined independently for each snapshot.
For each snapshot, you have two options:
* **Calculate Grid Center Auto** — Click this button and the platform computes the docking box center automatically based on the current protein structure. Work through each snapshot one by one.
* **Calculate Grid Center via Residue** — Specify your known binding site residues by name to anchor the box. This is faster if you have a well-characterized pocket and ensures the box tracks the same region across all conformations.
After calculating the grid center for each snapshot, confirm the selection and verify the box is correctly centered on your pocket of interest before proceeding to the next. The protein structure will look different in each snapshot as it reflects a different point in the MD trajectory.
Repeat this process for every snapshot in your ensemble.
Navigate to **Data Engineering** within the ensemble docking interface, and drag-drop the CSV file containing your compound library.
Click **Run Pipeline**. The docking algorithm will send every compound in your library through each protein snapshot — Trajectory 1, 2, 3, 4, 5, and so on — generating a **complete matrix of docking data** across all conformational states.
***
## Understanding the Ensemble Docking Output
The output is a multi-dimensional dataset where each compound has a docking score against every protein snapshot. This data matrix enables several levels of analysis:
**Consistent binders:** Compounds that score well across most or all snapshots are likely robust, conformationally-insensitive binders. These are your highest-confidence leads for progression.
**Conformation-selective binders:** Compounds that score strongly against only specific snapshots may act as conformational selectors or allosteric modulators, stabilizing particular protein states. These can be highly valuable for targeted biology.
**False positive filtration:** A compound that scores well in one snapshot but poorly across the rest is likely a docking artifact of that specific geometry rather than a true binder. Ensemble docking dramatically reduces this class of false positives compared to single-structure docking.
Analyze your full results matrix in the **RevAnalytics** module.
***
## Next Steps
After ensemble docking, your prioritized compound list is ready for:
* [**RevFEP**](/docs/revfep) — Alchemical free energy perturbation calculations for rigorously ranking your top candidates by binding affinity
* [**RevMD-Bind**](/docs/revmd-bind) — Full protein-ligand MD simulations to characterize binding stability and residence time for shortlisted leads
* **RevAnalytics** — Deep-dive statistical analysis of your ensemble docking score matrix, interaction fingerprinting, and pose clustering
* [**Static & Flexible Docking**](/tutorials/rev-bind/static-flexible-docking) — Review the earlier stages of the docking workflow if you are revisiting this tutorial
# Production VS with 2M+ Compounds
Source: https://docs.revilico.bio/tutorials/rev-bind/production-vs
A step-by-step walkthrough of production-scale high-throughput virtual screening using RevDock and RevScreen — from batch configuration to parallel pipeline execution
## Overview
This tutorial covers production-scale high-throughput virtual screening (HTS) in **RevBind's RevDock/RevScreen engine**. Unlike small-batch static, flexible, or ensemble docking runs, this workflow is designed for screening **multi-million compound libraries** efficiently using GPU-enabled parallel batch execution.
[High Throughput Screening with RevDock and RevScreen — Watch Video](https://www.loom.com/share/ee8e72f49d1c4fc08769f08496551dd8)
***
## When to Use This Workflow
This workflow is for **production high-throughput screening** — not the same as running static, flexible, or ensemble docking on a small compound set. Use it when:
* You have a large compound library (hundreds of thousands to millions of compounds)
* You want to maximize GPU throughput by running multiple pipelines in parallel
* You are executing a primary screen before downstream hit filtering and validation
For smaller sets or higher-accuracy modes, see [RevScreen - Static & Flexible Docking](/tutorials/rev-bind/static-flexible-docking) and [RevScreen - Ensemble Docking](/tutorials/rev-bind/ensemble-docking).
***
## Key Parameters
| Parameter | Recommended Value | Notes |
| ------------------- | ---------------------- | --------------------------------------------------- |
| **Batch size** | \~150,000 compounds | Optimized for GPU-enabled throughput |
| **Exhaustiveness** | 8 | Do not increase — compute scales disproportionately |
| **Parallelization** | One pipeline per batch | Run all batches simultaneously |
***
## Step-by-Step Walkthrough
From the Revilico OS dashboard, open **Revbind**, then navigate to **RevDock → RevScreen**. This is the production HTS interface — distinct from the static, flexible, and ensemble docking modes.
Load the protein structure you want to screen against. Confirm the correct structure is selected before proceeding — for example, the apo (no inhibitor) form of your target receptor.
If you haven't yet characterized the binding site, run a **RevPocket** analysis first to identify the druggable pocket and obtain the coordinates you'll need for box placement.
Review the target site and define the screening box:
* Determine where the box should be placed based on prior structural analysis
* Identify the key amino acids in the target binding site
* Calculate the **grid center** for the docking region
* Adjust box position until it is centered correctly on the pocket
For production runs, box placement is typically based on prior RevPocket output or guidance from your computational team.
Set **Exhaustiveness = 8** for high-throughput mode. Do not increase this value — compute cost scales disproportionately at higher settings. The goal is efficient throughput with usable screening results, not maximum conformational sampling.
Raising exhaustiveness above 8 for a 2M+ compound screen will cause a significant and non-linear increase in compute time and cost. Reserve higher exhaustiveness values for smaller, late-stage confirmation runs.
For large libraries, you will receive compounds pre-split into batches of approximately **150,000 compounds each**. For a 2 million compound library, this generates roughly 13–14 batches.
If your library is not already batched, you can:
* Ask the Revilico team to split it for you at no cost
* Use Claude Code or RevAgent to batch it yourself
Each batch should be a separate file (e.g., `Batch_1.sdf`, `Batch_2.sdf`, etc.).
Different compound libraries (e.g., Enamine REAL, custom sets) can be screened separately using the same workflow. Each library gets its own set of batch pipelines.
Select **Batch 1** as your compound library. Confirm:
* The correct **compound library batch** is loaded
* The correct **protein structure** is selected
* The **box and grid settings** are confirmed from Step 3
Click **Run Pipeline**. Once the pipeline is created successfully, Batch 1 begins processing.
This is the key step that makes production HTS efficient. Rather than waiting for each batch to complete sequentially, run all batches **in parallel**:
1. Clear Batch 1 from the configuration
2. Load **Batch 2**, confirm settings, click **Run Pipeline**
3. Repeat immediately for Batch 3, Batch 4, and so on
Each batch gets its own pipeline. All pipelines run simultaneously across the GPU cluster, compressing total wall-clock time dramatically compared to a single sequential job.
The box placement, protein structure, and exhaustiveness settings remain identical across all batches. The only thing that changes is the compound library file.
Once all pipelines complete, your screening data becomes available for download. Export the results and run filtration and analytics using your preferred AI agent or analysis workflow.
A follow-up tutorial will cover the full post-screen analytics step in detail.
***
## Parallelization Strategy
The core performance principle of this workflow is **parallelization over batching**:
* A single 2M compound job run sequentially is slow and resource-inefficient
* Splitting into \~150K batches and running all pipelines simultaneously compresses wall-clock time significantly
* GPU utilization stays high across the cluster rather than being bottlenecked by a single large job
This is why batch size (\~150K) is tuned to GPU throughput capacity, and exhaustiveness is kept at 8 — the combination maximizes screening velocity without sacrificing result quality.
***
## Next Steps
After your HTS run completes, the standard hit progression workflow is:
* **RevAnalytics** — Filter by docking score, select top-ranked compounds, analyze interaction fingerprints
* [**Flexible Docking**](/tutorials/rev-bind/static-flexible-docking) — Re-dock your top hits with flexible residues and CNN rescoring for higher-confidence rankings
* [**Ensemble Docking**](/tutorials/rev-bind/ensemble-docking) — Run your highest-priority compounds against MD-sampled protein conformations for maximum accuracy
* [**RevFEP**](/docs/revfep) — Compute rigorous binding free energies for your final shortlist
# RevScreen - Static & Flexible Docking
Source: https://docs.revilico.bio/tutorials/rev-bind/static-flexible-docking
A step-by-step walkthrough of Rigid Receptor and Flexible Docking in the RevBind Virtual Screening Engine
## Overview
This tutorial covers the first two docking modes available in **Rev-Bind's Virtual Screening Engine**: Rigid Receptor Docking and Flexible Docking. Together, they form the foundation of Revilico's tiered virtual screening workflow — from rapid, high-throughput screening all the way to structure-activity-guided hit refinement.
[Static Docking and Flexible Docking Review — Watch Video](https://www.loom.com/share/3b2c9a52e0c5473fbd6e8ebbdd496646)
***
## Docking Modes at a Glance
Rev-Bind offers three docking modes, each designed for a different stage of your screening campaign:
| Mode | Protein Flexibility | Rescoring | Best For |
| -------------------- | -------------------- | ------------- | ---------------------------------- |
| **Rigid Receptor** | Static | Standard | Rapid library-scale screening |
| **Flexible Docking** | Key residues flex | CNN rescoring | SAR-guided hit refinement |
| **Ensemble Docking** | Full MD-sampled flex | Standard | High-accuracy late-stage screening |
This tutorial covers **Rigid Receptor** and **Flexible Docking**. For Ensemble Docking, see the [Ensemble Docking tutorial](/tutorials/rev-bind/ensemble-docking).
***
## Part 1: Rigid Receptor Docking
Rigid receptor docking keeps the protein structure static and samples multiple conformations of each ligand within the binding pocket. This is your fastest, highest-throughput screen — ideal for processing large compound libraries quickly.
### How It Works
The algorithm holds the protein fixed, then exhaustively samples ligand conformations inside the defined binding box. Each pose is scored using an energy-based function, and the top-ranked conformations are returned for downstream analysis.
### Step-by-Step Walkthrough
From the Revilico OS dashboard, go to **Binding Chemistry → Virtual Screening Engine**. You will see three docking modes: Rigid Receptor, Flexible, and Ensemble. Select **Rigid Receptor Docking**.
Enter a descriptive name for your screening campaign. This name will identify the pipeline in your results history.
Upload your compound library (SDF or SMILES format), or select demo files to follow along. In this walkthrough, we use **800 AXL inhibitor compounds**.
Search for and select your target protein. In this walkthrough, we use **AXL (no inhibitor)** — the apo form of the receptor. The platform will automatically populate the 3D protein structure.
Using structural information from literature or a prior **RevPocket** search, draw the docking box around your binding pocket of interest. This tells the algorithm where on the protein to compute interaction energies.
Run a RevPocket analysis first if you have not yet characterized the binding site. It will give you the coordinates and key residue information you need to set the box accurately.
Adjust the **exhaustiveness** parameter to control the depth of conformational sampling. Higher values increase accuracy at the cost of compute time. See the [RevScreen documentation](/docs/revscreen) for guidance on calibrating this parameter to your campaign needs.
Click **Run Pipeline → Create → Close**. Your rigid receptor docking campaign is now queued. Results will populate in the RevAnalytics pane when the job completes.
***
## Part 2: Flexible Docking
Flexible docking builds on rigid receptor docking by allowing selected binding site residues to move during the simulation. It also applies **Convolutional Neural Network (CNN) rescoring** to improve pose quality assessment — delivering more confident binding predictions when you already have SAR or structural context.
### When to Use Flexible Docking
Use flexible docking as a **secondary filter** after rigid receptor screening, particularly when:
* You have SAR data from previous campaigns or literature identifying key binding interactions
* You know which residues engage with your ligands (e.g., hinge region residues for kinases)
* You want CNN rescoring to re-rank your top poses with greater accuracy
### How CNN Rescoring Works
After docking, the Convolutional Neural Network rescoring layer re-evaluates each pose using a 3D-aware deep learning model trained on crystallographic binding data. This step helps correct for cases where classical docking scoring functions misrank poses — especially useful for novel chemotypes or compounds with unusual binding geometries.
### Step-by-Step Walkthrough
From the Virtual Screening Engine, select **Flexible Docking**.
Enter a pipeline name. If you are continuing from a rigid receptor run, your compound library and protein will carry over automatically. Otherwise, re-upload your compound library and re-select your protein.
This is the critical step that differentiates flexible from rigid docking. Using your SAR knowledge or structural literature, identify which binding site residues you want to allow to flex during the calculation.
In the walkthrough example, we select **Pro621** and additional hinge region residues based on known kinase binding interactions. Click on residues in the 3D viewer to toggle their flexibility.
Not all residues can be set to flex — the platform will indicate which ones are eligible. Focus on residues for which you have structural evidence of involvement in key binding interactions.
Position the docking box around your binding pocket, just as in rigid receptor docking. Calibrate the exhaustiveness parameter based on your campaign requirements.
Click **Run Pipeline → Create**. The algorithm will sample ligand conformations, allow the designated residues to flex, and apply CNN rescoring to refine pose rankings. Your results will appear in RevAnalytics when complete.
***
## Comparing Results
After running both screens, compare the rank order from rigid vs. flexible docking. Compounds that rank highly in both are your highest-confidence hits for progression. Discordant rankings — especially where a compound improves under flexible docking — often indicate that the flexible residues you selected are critical for binding.
***
## Next Steps
Once you have completed rigid and flexible docking screens, your top-ranked compounds are ready for:
* [**Ensemble Docking**](/tutorials/rev-bind/ensemble-docking) — For highest-priority compounds, run ensemble docking against MD-sampled protein conformations for maximum accuracy
* **RevAnalytics** — Analyze docking scores, interaction fingerprints, and binding pose quality across your screened library
* [**RevFEP**](/docs/revfep) — Run rigorous free energy perturbation calculations on your shortlisted candidates
* [**RevMD-Bind**](/docs/revmd-bind) — Validate top compounds with full protein-ligand MD simulations
# AlphaFold
Source: https://docs.revilico.bio/tutorials/rev-target/alphafold
Predict the 3D structure of a protein from its amino acid sequence using Revilico's AlphaFold engine
## Overview
**AlphaFold** is Revilico's protein structure prediction engine. It takes a raw amino acid sequence as input and generates the full 3D structure of the protein — enabling downstream analyses like pocket identification, docking, and MD simulation without requiring an experimental crystal structure.
This is the starting point for any target-based drug discovery campaign where no experimental structure is available.
[Exploring the AlphaFold Workflow for Protein Structure Prediction — Watch Video](https://www.loom.com/share/35a1d1492f3d4a7e8e664546f09febc9)
***
## How AlphaFold Works
AlphaFold predicts a protein's 3D structure directly from sequence. The core breakthrough is that it solves the protein folding problem — mapping a linear amino acid sequence to the precise three-dimensional geometry that determines the protein's function and druggability.
Every engine in Revilico includes documentation on the right-hand side of the interface:
* **Documentation** — an overview of what the engine does and its scientific context
* **Configuration** — a step-by-step guide through the input parameters
* **Revilico Guide** — AI-powered assistant that searches through Revilico documentation and suggests answers to your questions
* **Interpreter** — reads your screen and helps interpret outputs across different engines
***
## Retrieving Your Protein Sequence
The AlphaFold engine requires a **single cohesive amino acid sequence** as input. The recommended source is **UniProt**.
Navigate to [UniProt](https://www.uniprot.org/) and search for the gene of interest. Select the human isoform where relevant.
On the UniProt entry page, scroll to the **Structure** section. If an AlphaFold structure already exists for your protein, you can download it directly — skipping the need to run a new prediction. Review the available variants (canonical, isoforms, etc.) and choose the one appropriate for your target.
If no pre-computed structure is available, go to the **Sequence** tab and copy the full amino acid sequence.
The engine requires the sequence in a **single unbroken line** with no spaces, line breaks, or irregular characters. If your sequence is formatted with breaks or whitespace, use the Revilico Interpreter to reformat it into a clean single-line string before pasting it in.
***
## Running an AlphaFold Pipeline
From the Revilico OS dashboard, open the AlphaFold engine.
Enter a descriptive pipeline name (e.g., `EGFR-alphafold-canonical`). This name will identify the run in your pipeline history.
Paste your amino acid sequence into the sequence input field. Ensure it is in a single cohesive line with no breaks or spaces.
If your sequence contains irregularities, paste it into the Revilico Interpreter and ask it to reformat the sequence into a single clean line.
Advanced parameters are pre-set to optimized defaults. Unless you have a specific reason to modify them, leave these as-is and proceed.
Click **Run Pipeline**. You will see a confirmation that the pipeline has been created and queued.
***
## Monitoring and Viewing Results
Once your pipeline is created, track it from the **Command Center** — the central hub for all pipeline activity.
### Checking Pipeline Status
Navigate to the Command Center from the top navigation. Your AlphaFold run will appear with its current status. You can monitor multiple pipelines simultaneously from this view.
### Viewing 3D Structure Output
When the run completes, open the results from the Command Center:
| Output | Description |
| ----------------------- | ----------------------------------------------------------------------------------- |
| **3D Structure Viewer** | Interactive visualization of the predicted protein structure |
| **Rank Number** | Model confidence ranking — lower rank numbers indicate higher-confidence structures |
| **3D Settings** | Adjust rendering, visibility, and coloring of the structural model |
| **Analytics** | Confidence metrics (pLDDT scores) and structural quality measures |
### Downloading the Structure
Click **Download** to export the predicted structure as a PDB file. This file can be used directly as input for downstream Revilico engines including RevPocket, RevScreen, and RevMD.
***
## Next Steps
With your AlphaFold structure in hand, the typical workflow continues:
* [**RevPocket**](/tutorials/rev-target/revpocket) — Identify druggable binding sites on your predicted structure before running docking
* [**RevScreen - Static & Flexible Docking**](/tutorials/rev-bind/static-flexible-docking) — Run virtual screening campaigns against your defined binding pocket
* [**RevScreen - Ensemble Docking**](/tutorials/rev-bind/ensemble-docking) — Use MD-sampled protein conformations for maximum-accuracy screening
* [**RevMD-Bind**](/docs/revmd-bind) — Validate binding stability of top candidates with protein-ligand MD simulation
# RevPocket
Source: https://docs.revilico.bio/tutorials/rev-target/revpocket
Identify and characterize druggable binding sites on protein structures using Revilico's pocket search engine
## Overview
**RevPocket** is Revilico's pocket search engine. It takes a protein structure as input and identifies all druggable binding sites on that protein — generating a ranked list of pockets with detailed chemistry profiles to guide downstream docking and screening campaigns.
RevPocket is typically run after generating or sourcing a protein structure, and before setting up a virtual screening campaign in RevScreen.
[Identifying Druggable Binding Sites in Protein Structures — Watch Video](https://www.loom.com/share/9c7f26770d2b446dadfea12f31550fee)
***
## How RevPocket Works
RevPocket takes a pure protein structure — a PDB file with no ligand or inhibitor present — and runs a cavity detection algorithm to identify all geometrically and chemically favorable binding sites. These pockets are ranked and characterized by their druggability, geometry, and chemical properties.
The results help you decide:
* **Which pocket to target** — the active site, an allosteric site, or a cryptic site
* **What chemistry to pursue** — polarity, hydrophobicity, and surface area drive compound design decisions
* **How to set up your docking run** — the pocket coordinates feed directly into RevScreen’s docking configuration
***
## Input Requirements
RevPocket requires a **PDB file of the apo protein** — the protein structure with no inhibitor, cofactor, or ligand bound. This ensures the pocket detection algorithm identifies natural cavities rather than pre-defined binding geometries from a co-crystal.
Your protein files are managed in the **File Manager** (Data Engineering) section of the Revilico OS dashboard.
Use the AlphaFold engine to generate a structure if you don’t yet have an experimental PDB file for your target. See the [AlphaFold tutorial](/tutorials/rev-target/alphafold) for the full workflow.
***
## Running a RevPocket Pipeline
From the Revilico OS dashboard, go to **RevTarget → RevPocket**.
Enter a name for this pocket search run (e.g., `AXL-apo-pocket-search`).
In the **PDB File Input** field, select your apo protein PDB file from the File Manager. Choose the file without any inhibitor present — the pure protein structure.
All parameters are pre-set to optimized defaults. If you’re an advanced user, click through to the documentation to understand each parameter in detail. For most use cases, the defaults are appropriate.
Click **Create PocketSearch**. You will see a confirmation: *PocketSearch pipeline created successfully.*
***
## Reading Your Results
Navigate to the **Analysis** section and load your completed pipeline to explore the pocket results.
### Pocket Rankings
RevPocket identifies and ranks up to 10 distinct pockets on your protein. Pockets are displayed as numbered overlays on the 3D structure. Select any individual pocket to isolate it in the 3D viewer and inspect its geometry.
### Pocket Analytics
Below the 3D viewer, a full analytics table provides chemical characterization for each pocket:
| Property | Description |
| ------------------------------------------ | -------------------------------------------------------------------------------------------------------- |
| **Druggability Score** | A composite metric estimating how likely this pocket is to bind a small molecule drug |
| **Solvent-Accessible Surface Area (SASA)** | The surface area exposed to solvent — larger values indicate more open, accessible pockets |
| **Polarity** | The ratio of polar residues lining the pocket — guides selection of polar vs. hydrophobic compounds |
| **Hydrophobicity** | The extent of hydrophobic character — high hydrophobicity pockets tend to bind lipophilic fragments well |
These properties together give you a deep chemical picture of each binding site before committing any compound library to a docking run.
### Interpreting Results in Biological Context
To decide which pocket to target, cross-reference the RevPocket output with:
* **Published literature** — review papers on known inhibitors or substrates for your target
* **Co-crystal structures** — if an inhibitor-bound structure exists, identify which RevPocket cavity corresponds to the bound inhibitor pocket
* **Substrate biology** — understand what molecule naturally engages this protein in its biological pathway, then target the pocket that substrate occupies
When a co-crystal structure or known inhibitor exists for your target, download it and visually overlay it with the RevPocket output. The pocket that contains the known ligand is almost always your primary target pocket.
***
## Next Steps
With your target pocket defined, you’re ready to move into screening:
* [**AlphaFold**](/tutorials/rev-target/alphafold) — If you still need a protein structure, generate one first
* [**RevScreen - Static & Flexible Docking**](/tutorials/rev-bind/static-flexible-docking) — Use your identified pocket coordinates to set up a virtual screening campaign
* [**RevScreen - Ensemble Docking**](/tutorials/rev-bind/ensemble-docking) — Run ensemble docking with MD-sampled conformations for higher-accuracy screening
* [**Boltz2 Cofolding**](/tutorials/rev-bind/boltz2-cofolding) — Use AI cofolding as an alternative for early-stage pose prediction without pocket pre-definition