Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 6 additions & 0 deletions CHANGES.txt
Original file line number Diff line number Diff line change
@@ -1,5 +1,11 @@
Changelog
---------
# unreleased
* mmCIF input read with biotite, as a single file (`-f complex.cif`) or as separate receptor and ligand files (`--receptor`, `--ligand`)
* chain options (`--chains`, `--peptides`, `--intra`, `--regions`) accept the original (multi-character) chain IDs of mmCIF files
* new `--npz` option for mmCIF input: interactions as numpy arrays with atom indices per input file
* new Python API `plip.profile.profile_system`

# 3.0.1
* remove setup.{py,cfg}, update pyproject.toml, use openbabel 3.2.0

Expand Down
52 changes: 52 additions & 0 deletions CONTEXT.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,52 @@
# PLIP

Profiling of non-covalent interactions between a receptor and its ligands in 3D structures.

## Language

**System**:
A single unit of analysis: one Receptor together with its Ligand(s), regardless of whether it is supplied as one structure file or as several.
_Avoid_: complex, structure (when meaning the whole input)

**Receptor**:
The protein part of a System, made of one or more chains.
_Avoid_: protein (when meaning the role), target

**Ligand**:
A molecule whose interactions with the Receptor are profiled. When supplied as its own file, the Ligand is the entire content of that file, whether it is a small molecule or a peptide; a lone metal ion is a Ligand as well.
_Avoid_: hetero group (when meaning the role)

**Surroundings**:
Everything in the System other than the Ligand currently being profiled — including the other Ligands — exactly as it would be present in a single structure file.

**Split input**:
A System supplied as several files, each explicitly declared as a Receptor file or a Ligand file.
_Avoid_: multi-file mode, PLINDER mode

**Single-file input**:
A System supplied as one file containing both Receptor and Ligand(s); Ligands are detected the same way as for PDB input.

**Chain mapping**:
The record of how chain identifiers from the source files were renamed when merging a Split input, so that each chain in the System is unique and traceable to its source file.

**Inter-chain interaction**:
An interaction between two chains of a System, one of which is treated as the Ligand (e.g. protein–protein or protein–peptide).
_Avoid_: PPI (when the partner may be a peptide)

**Intra-chain interaction**:
An interaction between residues of the same chain.

## Relationships

- A **System** has exactly one **Receptor** and one or more **Ligands**
- A **Split input** yields exactly one **System** and exactly one **Chain mapping**
- Every chain in a **System** traces back to exactly one source file via the **Chain mapping**
- Each **Ligand** is profiled separately, against the **Receptor**, with the other **Ligands** present in the **Surroundings**
- A **Split input** and a **Single-file input** of the same structure describe the same **System** and must yield the same interactions
- In a **Split input**, each Ligand file is exactly one **Ligand** and is never discarded as an artifact, buffer or crystallisation additive; in a **Single-file input**, Ligands are detected and filtered as for PDB
- A **System** with no **Ligands** is still valid: it yields no ligand interactions but can be profiled for **Inter-chain** and **Intra-chain interactions**
- Users always name chains by their original identifiers; the **Chain mapping** is internal and only surfaces when an identifier is ambiguous

## Flagged ambiguities

- "structure file" vs "System": one System may come from many files; resolved — the System is the unit, files are only its sources.
53 changes: 53 additions & 0 deletions DOCUMENTATION.md
Original file line number Diff line number Diff line change
Expand Up @@ -118,6 +118,59 @@ Please note that detection within a chain takes much longer than detection of pr
### Interactions of Molecules with DNA/RNA
PLIP can characterize interactions between ligands and DNA/RNA. A special mode allows to switch from treating DNA/RNA molecules as ligands to treating them as part of the receptor in the structure. If a protein is present, too, interactions of the ligand with both, protein and nucleic acids, will be shown. To use this mode, start PLIP with the option `--dnareceptor`.

## Analyze mmCIF structures with PLIP
mmCIF files (`.cif`, `.cif.gz`, `.mmcif`) are read with [biotite](https://www.biotite-python.org) and then analyzed by exactly the same pipeline as PDB files (see `docs/adr/0001-mmcif-read-with-biotite-routed-through-pdb-text.md`). The terms used below are defined in `CONTEXT.md`.

### Single-file and Split input
A System can be given as one mmCIF file with receptor and ligands (Single-file input), in which case ligands are detected as for PDB files:

```bash
$ plip -f complex.cif -xt --npz
```

or as separate receptor and ligand files (Split input), e.g. from PLINDER:

```bash
$ plip --receptor receptor_A.cif receptor_B.cif --ligand ligand_1.cif ligand_2.cif -xt --npz
```

In a Split input, each ligand file is exactly one ligand, consisting of all its atoms, and is never discarded as artifact, buffer or ion. A ligand file with a single metal ion is a ligand of type `ION`, as for PDB input. All other ligands are also present in the surroundings while one ligand is profiled. `--ligand` is optional, e.g. for interactions between receptor chains only.

Chains and residues are identified by the `auth_*` fields (`label_*` only if `auth_*` is missing). Internally, every chain is given a unique single-character ID and residue names longer than three characters (e.g. `GLC-F`) an alias; reports always show the original identifiers and contain the chain mapping. `--chains`, `--peptides`/`--inter`, `--intra` and `--regions` take the original chain IDs, e.g. `--chains "[[1.A], [1.B]]"`. If a chain ID occurs in more than one file, it is ambiguous and PLIP stops with an error showing the chain mapping. Atom indices in XML and TXT reports (`DONORIDX`, `LIG_IDX_LIST`, ...) refer to the internal merged structure; use the NPZ output for indices into the input files. PyMOL visualization (`-p`, `-y`) and `--stdout` are not supported for mmCIF input.

### NPZ output
With `--npz` (mmCIF input only), PLIP writes all interactions of the System as numpy arrays (`<name>.npz`):

| Array | Shape | Content |
|---|---|---|
| `pairs` | (P, 6) | `[rec_file, rec_atom, lig_file, lig_atom, type, interaction_id]`, one row per atom pair. Interactions between atom groups (rings, charged groups) contribute all pairs of their atoms. For metal complexes, the metal is the ligand-side atom and the coordinated atom the receptor-side atom (see `location`) |
| `rec_chain`, `lig_chain` | (P,) | original chain IDs of both atoms of a pair |
| `files`, `file_role` | (F,) | input files (`rec_file`/`lig_file` index into them) and their role (`complex`, `receptor`, `ligand`) |
| `interaction_types` | (8,) | names of the codes in `type`: hydrophobic, hbond, waterbridge, saltbridge, pistacking, pication, halogen, metal |
| `interaction_type`, `interaction_ligand` | (I,) | type and ligand (index into `ligand_*`) of each interaction |
| `resnr`, `restype`, `reschain`, `resnr_lig`, `restype_lig`, `reschain_lig` | (I,) | residues of the interaction partners |
| `dist`, `centdist`, `dist_h_a`, `dist_d_a`, `dist_a_w`, `dist_d_w`, `don_angle`, `acc_angle`, `water_angle`, `angle`, `offset`, `rms` | (I,) | geometry, NaN if not applicable |
| `sidechain`, `protisdon`, `protispos`, `protcharged`, `coordination`, `complexnum`, `water_file`, `water_atom` | (I,) | integer features, -1 if not applicable |
| `donortype`, `acceptortype`, `stacking_type`, `lig_group`, `metal_type`, `target_type`, `location`, `geometry` | (I,) | string features, empty if not applicable |
| `ligcoo`, `protcoo`, `watercoo` | (I, 3) | coordinates as in the XML report |
| `ligand_hetid`, `ligand_chain`, `ligand_resnr`, `ligand_longname`, `ligand_type`, `ligand_file`, `ligand_smiles`, `ligand_inchikey` | (L,) | ligands (`ligand_file` is -1 if the ligand is not a ligand file) |
| `chain_mapping_file`, `chain_mapping_original`, `chain_mapping_internal` | (C,) | chain mapping |

An atom index is the position of the atom in the `_atom_site` loop of its file within the selected model, i.e. the index into the AtomArray returned by `biotite.structure.io.pdbx.get_structure(pdbx.CIFFile.read(path), model=1, altloc="all")`. Hydrogens are never referenced; polar hydrogens are added by OpenBabel and interactions refer to their heavy atoms.

### Python API
The same analysis is available as a function, which returns the NPZ content as a dict of numpy arrays:

```python
from plip.profile import profile_system

arrays = profile_system(receptor=['receptor.cif'], ligands=['ligand_1.cif', 'ligand_2.cif'],
npz='system.npz', xml='system.xml')
arrays = profile_system(structure='complex.cif', chains=[['1.A'], ['1.B']])
```

Options correspond to the command line options (`peptides`, `intra`, `regions`, `model`, `keepmod`, `nohydro`, ...). PLIP keeps its settings in the global module `plip.basic.config`; `profile_system` restores them afterwards, but it is not thread-safe. Use processes for parallel runs.

## Changing detection thresholds
The PLIP algorithm uses a rule-based detection to report non-covalent interaction between proteins and their partners. The current settings are based on literature values and have been refined based on extensive testing with independent cases from mainly crystallography journals, covering a broad range of structure resolutions. For most users, it is not recommended to change the standard settings. However, there may be cases where changing detection thresholds is advisable (e.g. sets with multiple very low resolution structures).

Expand Down
14 changes: 14 additions & 0 deletions docs/adr/0001-mmcif-read-with-biotite-routed-through-pdb-text.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,14 @@
# mmCIF is read with biotite and routed through PDB text into the existing OpenBabel pipeline

mmCIF input (a single file, or a Split input of receptor and ligand files) is read and merged with biotite. The result is written as temporary PDB text, with chain IDs renamed to single characters and residue names longer than three characters given 3-character aliases. That text goes through the existing `PDBParser` and pybel path, and OpenBabel does all of the chemistry, exactly as it does for PDB input. Reports translate the aliases back to the original chain IDs and residue names.

We rejected three alternatives:
- **Reading mmCIF with OpenBabel directly.** Tested on PLINDER files, the reader drops chain IDs (it re-letters them A, B, …), sets the HETATM flag to false everywhere, sets the side-chain flag on every atom, adds no implicit hydrogens (so there is no aromaticity and no polar H), ignores `_chem_comp_bond` and charges, and silently fails on single-atom files.
- **Building an OBMol directly from biotite arrays.** This would keep bond orders from the file, but it means re-implementing what OpenBabel's PDB reader and PLIP's `PDBParser` do, which risks subtle divergence from PDB results.
- **Rewriting the core on biotite.** This amounts to a rewrite of PLIP.

## Consequences

- Bond orders from `_chem_comp_bond` are discarded and OpenBabel perceives them again, the same as for PDB input.
- A System can hold at most 62 chains, the number of single characters available for a PDB chain ID.
- A CIF file and a PDB file of the same structure yield identical interactions, and PDB input serves as the test oracle.
1 change: 1 addition & 0 deletions plip/basic/config.py
Original file line number Diff line number Diff line change
Expand Up @@ -35,6 +35,7 @@
CHAINS = None # Define chains for protein-protein interaction detection
REGIONS = None
COMPRESS = False # Compress XML and TXT report files
NPZ = False # Write interaction arrays in NPZ format (mmCIF input only)


# Configuration file for Protein-Ligand Interaction Profiler (PLIP)
Expand Down
Loading