> ## Documentation Index
> Fetch the complete documentation index at: https://docs.revilico.bio/llms.txt
> Use this file to discover all available pages before exploring further.

# RevHBond

> Hydrogen-bond acceptor (pKBHX) and donor (pKα) strength for every site in a molecule, from a calibrated DFT electrostatic potential

## Why Use This Engine?

RevHBond predicts how strongly each part of a molecule forms hydrogen bonds. For every acceptor atom (a carbonyl oxygen, a pyridine nitrogen, an ether oxygen) it returns a **pKBHX**, and for every O–H and N–H hydrogen it returns a **pKα**. It also combines the sites into one acceptor value and one donor value for the whole molecule.

Both numbers are log scales, so one unit is a tenfold stronger hydrogen bond. For reference, pyridine's nitrogen accepts at about 1.9 and DMSO's oxygen at about 2.5.

Use RevHBond when you need to know which atom does the work, not just how many donors and acceptors a molecule has. Typical uses:

* **Tune permeability and efflux.** Find the strongest acceptors and donors in a series, and check whether a change weakens the one that matters.
* **Compare isosteres.** Rank replacement groups by how strongly they accept or donate, on the same scale.
* **Rationalize binding.** Check whether an acceptor or donor in a ligand is strong enough to make a key contact with the protein.
* **Rank a series.** Sort molecules by overall acceptor or donor strength.

You can submit up to 25 neutral molecules in one run, from SMILES or from 3D structure files.

<Frame>
  <img src="https://mintcdn.com/revilicoinc/IRxho5RindUWmw9-/images/revhbondworkflow.png?fit=max&auto=format&n=IRxho5RindUWmw9-&q=85&s=cb826dd09ba8acf3ba8d4510c6a907ee" alt="RevHBond Workflow" width="3400" height="1880" data-path="images/revhbondworkflow.png" />
</Frame>

## Background

A hydrogen bond links a hydrogen on an electronegative atom (the **donor**) to a lone pair on another atom (the **acceptor**). Its strength is usually measured as the equilibrium constant for forming a 1:1 complex with a reference partner. pKBHX is the standard acceptor scale: the base-10 logarithm of the binding constant with 4-fluorophenol in carbon tetrachloride. pKα is the matching scale for donors; it is not the acid dissociation constant pKa.

$$
\text{pK}_{\text{BHX}} = \log_{10} K_{\text{BHX}}
$$

Because both scales are logarithms, a site at 2.0 binds ten times more strongly than one at 1.0.

RevHBond predicts these values from the molecule's **electrostatic potential**, the energy a positive test charge would feel at each point around the molecule. A lone pair shows up as a minimum in the potential a little over an ångström in front of the acceptor atom. The deeper that minimum, the stronger the acceptor. For a donor, the potential is read at a probe point just past the hydrogen.

The engine computes the potential once per molecule with density functional theory, reads it at each acceptor minimum and donor probe point, and converts each reading to a strength with a fitted line:

$$
\text{pK} = a \, V + b
$$

where $V$ is the potential at the site, and $a$ and $b$ are fitted to measured values. Acceptors are split into 13 types, each with its own line (see [Validation](#validation)). Every O–H and N–H donor uses one shared line.

An atom with two lone pairs, such as a carbonyl oxygen, has two minima. RevHBond reports each one, plus a value for the atom that combines them. The atom value is the number an experiment on that atom measures.

The molecule-level value combines every site stronger than −1, because binding constants add:

$$
\text{pK}_{\text{molecule}} = \log_{10} \sum_{i} 10^{\,\text{pK}_i}
$$

A molecule with several good acceptors is therefore a stronger acceptor overall than any one of its sites.

**Level of theory.** The potential is always computed at r2SCAN-3c, on a geometry relaxed with the AIMNet2 machine-learned potential. RevHBond writes this as `r2SCAN-3c//AIMNet2`: potential on the left, geometry on the right. You can't change the method. The calibration lines were fitted to potentials from this protocol only, and a potential from another method would turn into a confidently wrong answer.

## Validation

RevHBond uses a published calibration (Wagen, ChemRxiv 2025, doi:10.26434/chemrxiv-2025-kv6d6-v2). Each acceptor line was fitted to measured pKBHX values. The table shows how many molecules each line rests on and how closely it reproduces them.

| Acceptor type | Examples | Molecules | MAE | RMSE |
| - | - | - | - | - |
| Amine | Triethylamine, piperidine, aniline | 171 | 0.212 | 0.324 |
| Aromatic N | Pyridine, imidazole, quinoline | 71 | 0.113 | 0.150 |
| Imine | Amidines, guanidines, oximes | 28 | 0.180 | 0.236 |
| Nitrile | Acetonitrile, benzonitrile | 28 | 0.144 | 0.198 |
| N-oxide | Pyridine N-oxide | 16 | 0.455 | 0.589 |
| Chalcogen oxide | DMSO, sulfones, sulfonamides | 17 | 0.186 | 0.224 |
| Pnictogen oxide | Phosphine oxides, phosphates | 16 | 0.437 | 0.549 |
| Carbonyl | Ketones, esters, amides, ureas | 128 | 0.160 | 0.208 |
| Ether/hydroxyl | THF, diethyl ether, alcohols | 99 | 0.188 | 0.239 |
| Thiocarbonyl | Thioureas, thioamides | 10 | 0.330 | 0.384 |
| Divalent S | Dimethyl sulfide, tetrahydrothiophene | 17 | 0.086 | 0.127 |
| Aromatic O | Furan, oxazole oxygen | 11 | 0.125 | 0.158 |
| Fluorine | Alkyl and aryl fluorides | 23 | 0.202 | 0.276 |

MAE and RMSE are in pKBHX units. Against measured values the method is typically within about 0.2 units. The worst cases are N-oxides and crowded amines: the calculation sees the amine's lone pair as strong, but a partner has trouble reaching it.

The donor line rests on 41 measured compounds, far fewer than the acceptor lines. Read pKα values as a ranking more than as exact values.

These are the fit statistics of the published calibration. RevHBond has not been benchmarked separately against an external reference.

## Running the Engine

Open **Quantum Chemistry** > **RevQuant** > **RevHBond**. The page has two tabs: **H-Bond Strength** to set up a run and **Analysis** to view results.

<Steps>
  <Step title="Name the run">
    Enter a **Name** (optional, up to 200 characters). Runs without a name are called `Hydrogen-bond strength` followed by the date.
  </Step>

  <Step title="Add your molecules">
    Drop files on **Molecules**, or click to browse. You can also drag a file from the **Data Engineering** panel. Accepted formats are `.sdf`, `.mol`, `.xyz`, `.pdb`, `.csv` and `.smi`. See [Preparing input files](#preparing-input-files).
  </Step>

  <Step title="Choose where the geometry comes from">
    Select a **Geometry** option. The default, **Find the lowest-energy shape**, is the right choice for a molecule on its own. The **Level of theory** strip shows the resulting protocol, such as `r2SCAN-3c//AIMNet2`. See [Geometry options](#geometry-options).
  </Step>

  <Step title="Adjust advanced settings (optional)">
    Open **Advanced Options** to set the number of **Starting conformers**, or the **Charge** for an XYZ file.
  </Step>

  <Step title="Set a runtime limit (optional)">
    Enter a **Runtime Limit (Credits)** to cap what the run can spend. Leave it blank to run until completion. See [Run time and credits](#run-time-and-credits).
  </Step>

  <Step title="Run the calculation">
    Click **Run calculation**, then **Confirm**. A notification shows how many molecules were submitted and the level of theory, and the page switches to the **Analysis** tab with your run selected.
  </Step>
</Steps>

### Inputs

| Setting | Default | Options and notes |
| - | - | - |
| Name | Optional | Up to 200 characters |
| Molecules | Required | `.sdf`, `.mol`, `.xyz`, `.pdb`, `.csv` or `.smi`; up to 25 MB per file, 25 molecules per run and 100 atoms per molecule, hydrogens included |
| Geometry | Find the lowest-energy shape | **Find the lowest-energy shape**, **Relax my conformation**, **Use my coordinates exactly**. See [Geometry options](#geometry-options) |
| Starting conformers | Sized from the molecule | 1 to 50. Left blank, it is ten plus five per rotatable bond, up to 50. Used only with **Find the lowest-energy shape** |
| Charge (XYZ only) | From structure | −10 to 10. Used only to work out the bonds in an XYZ file |
| Runtime Limit (Credits) | No limit | Optional whole number of credits. The run stops when it reaches the limit |

You can't choose the method or basis set. The potential is always computed at r2SCAN-3c, because the calibration is valid only for that protocol.

### Geometry options

The electrostatic potential depends on the molecule's shape, so the **Geometry** setting decides which shape the strengths describe.

| Option | Level of theory | What it does | Use for |
| - | - | - | - |
| **Find the lowest-energy shape** | `r2SCAN-3c//AIMNet2` | Generates conformers, relaxes the best few with AIMNet2 and predicts on the lowest-energy one | A molecule on its own. This is the protocol the calibration was fitted on |
| **Relax my conformation** | `r2SCAN-3c//AIMNet2` | Keeps the shape you uploaded and only relaxes it with AIMNet2 | A bound pose, where an internal hydrogen bond present in the pose should stay |
| **Use my coordinates exactly** | `r2SCAN-3c//submitted geometry` | Predicts on the coordinates as uploaded, with no optimization | A structure already optimized elsewhere. Anything rougher moves the results |

<Note>
  **Relax my conformation** shows a warning before you submit: without a conformer search, the submitted conformation is kept. For a molecule on its own, the calibration assumes the lowest-energy conformer.
</Note>

### Preparing input files

**Submit the neutral form of each molecule**: the free base, not the salt. The calibration covers neutral molecules only. A charged molecule fails with that reason rather than returning numbers that look plausible but are wrong.

**CSV or SMI (SMILES).** SMILES carry no 3D shape, so one is generated with a force field. Use **Find the lowest-energy shape** for SMILES input. **Use my coordinates exactly** isn't available, because there are no coordinates to keep. With **Relax my conformation**, the form warns you that the force-field shape is the starting point.

**SDF.** An SDF can hold many molecules, one per record.

**MOL, XYZ or PDB.** One molecule per file. You can select several of these files at once, up to 25. To submit many molecules in other formats, use a multi-record SDF or a CSV instead.

**XYZ.** An XYZ file states no bonds or charge. If the molecule isn't neutral, set **Charge** under **Advanced Options** so the bonds can be worked out. Every other format carries its own charge.

Atom numbers in the results follow the order of atoms in your input file.

### Run time and credits

Each molecule needs a conformer search and one DFT calculation. Expect about a minute for a small molecule and about ten minutes for a drug-sized one. A run processes one molecule at a time, so the total time grows with the number of molecules. Finished molecules are saved as they complete.

RevHBond bills 1 credit per minute of runtime while the calculation runs. You need enough credits to start a run.

* Leave **Runtime Limit (Credits)** blank to run until the calculation finishes, with no cap.
* Set a limit to cap the spend. When the run reaches it, the run stops with the status `terminated_budget_exceeded`.

A run also stops if your credit balance runs out.

## Viewing Results

Open the **Analysis** tab and click a run in the list. Click **Change Pipeline** to go back to the list. A run that is still in progress shows its percentage complete and refreshes automatically every 10 seconds. A run that failed or was stopped shows the reason.

### Run statuses

| Status | Meaning |
| - | - |
| `not_started` | Submitted and waiting for compute |
| `started` | Running |
| `processed` | Finished. Some molecules may still have failed; check the header |
| `error` | The run failed |
| `terminated_insufficient_credits` | Stopped because your credit balance ran out |
| `terminated_budget_exceeded` | Stopped because the run reached its **Runtime Limit (Credits)** |

The run header shows the level of theory, how many molecules completed out of the total, how many failed, and the total runtime.

### Molecule table

The table lists one row per molecule. Click a column heading to sort by it. Click a row to open that molecule's detail above the table; the first completed molecule opens automatically.

| Column | Shows |
| - | - |
| **#** | Position in the input |
| **Name**, **Formula** | Molecule name and formula |
| **Accepts (pKBHX)** | Molecule-level acceptor strength |
| **Strongest acceptor**, **Type**, **pKBHX** | The strongest acceptor atom, its acceptor type and its value |
| **Donates (pKα)** | Molecule-level donor strength |
| **Strongest donor**, **pKα** | The strongest donor hydrogen and its value |
| **Status** | `completed`, `submitted geometry` for a molecule predicted on its uploaded coordinates, or the error message for a failed molecule |

Values are shown to two decimals, since the calibration is good to about 0.2 units. If any molecule failed, turn on **Failures only** to list just those.

### Structure viewer

The **Structure** panel shows the molecule in 3D with every hydrogen-bond site drawn where it sits.

* **Acceptor sites** are spheres at each potential minimum, in front of the lone pair, joined to their atom by a dashed stick. The stick shows the direction a hydrogen bond would come in from.
* **Donor sites** are blue spheres at the probe point just past the hydrogen.
* Larger spheres are stronger sites. Labels give the atom and its value.

Acceptor colors follow fixed strength bands, the same on every molecule:

| Band | pKBHX or pKα |
| - | - |
| Very strong | 3 or more |
| Strong | 2 to 3 |
| Moderate | 1 to 2 |
| Weak | 0 to 1 |
| Very weak | Below 0 |

The bands are a reading aid, not part of the calculation.

Use **Both**, **Acceptors** or **Donors** to choose which sites to draw. Sites weaker than −1 are hidden by default; click **Hiding sites below −1** to show them. Click a site in the **Sites** tab to highlight it in the viewer.

Switch between **Computed geometry**, the shape the sites were calculated on, and **As submitted**, your uploaded structure. **As submitted** is available only when the geometry was optimized or searched, and it shows no sites. Click **SDF** or **XYZ** to download the structure in view.

### Sites tab

Two headline values give the molecule-level **Accepts (pKBHX)** and **Donates (pKα)**, each with its strength band. A molecule with no site above −1 shows "no site above −1".

**Acceptors** lists every acceptor atom of a calibrated type:

| Column | Shows |
| - | - |
| **Atom** | Element and atom number, following your input file |
| **Type** | The acceptor type, which decides which calibration line scores the atom. Hover to see how many measured molecules the line was fitted to and its average error |
| **pKBHX** | The atom's combined value |
| **Per lone pair** | The value at each potential minimum on the atom |

**Donors** lists every O–H and N–H hydrogen, with the hydrogen's atom number (**Hydrogen**), the atom it's bonded to (**On**) and its **pKα**.

Click **Sites CSV** to download every acceptor minimum and donor hydrogen for this molecule.

### Molecule tab

Shows any warnings for the molecule and how the numbers were made:

| Row | Shows |
| - | - |
| **Level of theory** | For example, `r2SCAN-3c//AIMNet2` |
| **Geometry** | Lowest-energy conformer found here, your conformation relaxed here, or as submitted |
| **Conformers** | Conformers generated → distinct → fully relaxed, for a conformer search |
| **Next-best conformer** | How much higher in energy the runner-up was, in kcal/mol |
| **Optimization converged** | Whether the AIMNet2 relaxation converged, and in how many steps |
| **Basis functions** | Size of the DFT calculation |
| **Time taken** | Runtime for this molecule, in seconds |

<Tip>
  If the next-best conformer is within about 1 kcal/mol, both shapes are populated at room temperature and their sites may differ. Treat the site values as describing the lowest conformer only.
</Tip>

### Report tab

The full plain-text calculation log for the molecule.

### Downloads

Click **Download** in the run header to get two files for the run. `{name}` is the run name, with spaces and special characters replaced by underscores.

* **`{name}-summary.csv`:** one row per molecule. Columns include `index`, `name`, `formula`, `smiles`, `molecular_pkbhx`, `strongest_acceptor_atom`, `strongest_acceptor_class`, `strongest_acceptor_pkbhx`, `molecular_pk_alpha`, `strongest_donor_atom`, `strongest_donor_pk_alpha`, `geometry_source`, `status` and `error`.
* **`{name}-results.json`:** the full batch document, with the run settings, counts and every site.

From a molecule's detail, you can also download:

| Button | File | Contents |
| - | - | - |
| **Sites CSV** | `molecule_{index}_sites.csv` | Every acceptor minimum and donor hydrogen for the molecule, one per row |
| **SDF** | `molecule_{index}_sites.sdf` | The computed geometry, with each site's pKBHX and pKα stored as atom properties |
| **XYZ** | `molecule_{index}_input_geometry.xyz` | The structure as you submitted it (from the **As submitted** view) |

## Limits

* Up to 25 molecules per run, 25 MB per file and 100 atoms per molecule, hydrogens included.
* Neutral molecules only. Submit the free base, not the salt.
* Elements are limited to those AIMNet2 covers: H, B, C, N, O, F, Si, P, S, Cl, As, Se, Br and I.
* The level of theory is fixed at `r2SCAN-3c` on an AIMNet2 geometry (or your submitted geometry).
* Only acceptors of the 13 calibrated types are scored. A molecule with none shows "No acceptor of a calibrated type in this molecule".
* Only O–H and N–H hydrogens are scored as donors. S–H is not scored and is left out of the molecule-level donor value.
* Sites weaker than −1 are left out of the molecule-level values.
* The donor calibration rests on 41 compounds; use pKα mainly to rank sites.
* Predictions are least reliable for N-oxides, pnictogen oxides and sterically crowded amines.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.