Clusters from pxr_diverse.csv (which lists one representative per
cluster); full membership taken from pxr_viewer_table.csv. MCS computed with
RDKit rdFMCS.FindMCS with ringMatchesRingOnly=True and
completeRingsOnly=False. Complete-rings-only is deliberately off: these clusters
pair a benzene with a saturated ring whose size varies between members, and demanding whole rings
discards the entire fused system — cluster 1 collapses to bare benzene.
The cluster representative column draws that member whole, with the MCS
highlighted orange so the shared core is seen in its ring context.
The caption gives the MCS size, the fraction of an average member it covers, and the MCS itself
as SMILES — written from the matching fragment of the representative, so it reads as a real
structure rather than a SMARTS query. Where the MCS covers only part of a ring those atoms appear
in that SMILES as an open chain: every cluster member has a ring there, the members just disagree
on its size. Exact queries are in the mcs_smarts column of
pxr_cluster_mcs.csv.
The fourth column is the distribution of pairwise RMSD over the MCS atoms
across every pair of cluster members (self-pairs excluded), measured on the crystallographic poses
in aligned_ligands.sdf. Those poses are already in a common frame from superposing the
proteins, so the cores are compared in place — nothing is re-superposed, and the value
is how differently the shared core sits in the site, not how differently it is shaped. Every
symmetry-equivalent mapping of the MCS is enumerated on both partners and the pair takes the best
one. Individual pairs are plotted over each box, since n runs from 3 to 28. All ten box plots
share the axis in this column header. Every value is in
pxr_cluster_mcs_rmsd_pairs.csv.
Those RMSDs are then clustered into binding modes: complete linkage on the
pairwise matrix, cut at 2 Å, so every pair inside a mode is within
2 Å of every other member rather than merely on average
— the usual threshold for calling two poses the same. The matrix column shows each cluster's
distances in dendrogram order on the shared colour scale, with the modes outlined; the diagonal is
blank because self-comparisons are excluded throughout. The same split is carried back onto the box
plot, where each pair is coloured by whether it falls
inside one mode or
across two — which is what the bimodal
distributions turn out to be. Assignments are in
pxr_cluster_binding_modes.csv.
The interaction barcode reads contacts straight from
interaction_fingerprint_implicit_H.csv — they are taken as given, not recomputed here. One lane per cluster
member, in the same dendrogram order and with the same short ids as the matrix beside it, so
a mode block there should show up as a band of similar barcodes here; one column per residue, on
the axis labelled in this header. Colour is the interaction: a residue making several takes the
most specific one, and van der Waals contacts draw as short pale stubs so the directional contacts
are what carries. That masking is real but small — 5 of the 233 contacted
residue-structure pairs make more than one directional interaction at once (H-bond donor
and acceptor, mostly). The residue axis is fixed across all ten barcodes: it is the
18 binding-site residues above plus every residue making a directional contact anywhere
in the set, 22 columns in all, so no H-bond or π-stack is hidden by the axis
choice. 5 van der Waals contacts at 4 residues outside it are the only
thing left out. Per-structure assignments are in pxr_cluster_interactions.csv.
The final column is the per-residue RMSF of the binding site across each
cluster's crystal structures, from aligned_structures/. The residue set is fixed
across all ten plots: a residue qualifies if any heavy atom comes within
5 Å of the ligand in at least 50% of the
45 clustered structures, and is kept only if present in all of them — giving
the 18 residues labelled in this header, compared over the same heavy atoms everywhere.
RMSF is the fluctuation about each cluster's own mean position; the structures already share a
frame, so nothing is superposed. Only the chain named in the viewer table is read: PXR crystallizes
as a dimer and only that copy was placed in the common frame. All ten bar plots share one
y-scale (0–2.8 Å), and any residue above
1 Å (the dashed line) is named on its bar. Values are in
pxr_cluster_pocket_rmsf.csv.
Tap any row to load that cluster's complexes — the page switches to the 3D & 2D structures tab as it loads, and the Table tab keeps its scroll position. There, drag the divider to trade width between the 2D grid on the left and the viewer on the right.
| Cluster | Size | Cluster representative | Pairwise MCS RMSD (Å)
same mode
different mode |
Binding modes |
Protein–ligand interactions
H-bond acceptor
H-bond donor
π-stacking
ionic / π-cation
van der Waals |
Binding-site RMSF (Å) |
|---|---|---|---|---|---|---|
| 1 | 8 |
OCNT-2395345 — MCS highlighted · 17 atoms,
71% of the average member CC=CCC(=O)N(CCC)c1ccccc1C |
median 2.27 Å ·
0.52–3.59 Å · 28 pairs
(7 same-mode) |
4 binding modes ·
largest 4/8 |
4.8 residues contacted per structure ·
21 directional · 0 residues
contacted by all 8 |
mean 0.73 Å ·
peak H407 1.82 Å |
| 2 | 7 |
OCNT-2316194 — MCS highlighted · 8 atoms,
37% of the average member CCC(=O)N(C)CC |
median 1.21 Å ·
0.30–3.37 Å · 21 pairs
(15 same-mode) |
2 binding modes ·
largest 6/7 |
4.1 residues contacted per structure ·
14 directional · 0 residues
contacted by all 7 |
mean 0.65 Å ·
peak H407 1.65 Å |
| 3 | 5 |
OCNT-2395728 — MCS highlighted · 21 atoms,
87% of the average member C=CC=CS(=O)(=O)NCC(O)c1cccc2ccccc12 |
median 6.20 Å ·
0.79–6.68 Å · 10 pairs
(2 same-mode) |
3 binding modes ·
largest 2/5 |
6.0 residues contacted per structure ·
18 directional · 0 residues
contacted by all 5 |
mean 0.63 Å ·
peak L209 1.71 Å |
| 4 | 5 |
OCNT-2308748 — MCS highlighted · 12 atoms,
48% of the average member CC=NCN1CCN(C=O)CC1 |
median 3.45 Å ·
1.23–6.16 Å · 10 pairs
(3 same-mode) |
3 binding modes ·
largest 3/5 |
5.4 residues contacted per structure ·
13 directional · 1 residue
contacted by all 5 |
mean 0.93 Å ·
peak L209 2.18 Å |
| 5 | 4 |
OCNT-2318616 — MCS highlighted · 21 atoms,
74% of the average member O=C(NCCS(=O)(=O)Nc1ccccc1)Nc1cc[nH]n1 |
median 0.94 Å ·
0.40–1.29 Å · 6 pairs
(6 same-mode) |
1 binding mode ·
largest 4/4 |
7.0 residues contacted per structure ·
14 directional · 3 residues
contacted by all 4 |
mean 0.41 Å ·
peak M323 0.97 Å |
| 6 | 4 |
OCNT-2316690 — MCS highlighted · 19 atoms,
83% of the average member Cc1cc(C(C)(C)C)cc(C)c1CS(=O)(=O)CC=N |
median 0.68 Å ·
0.35–0.98 Å · 6 pairs
(6 same-mode) |
1 binding mode ·
largest 4/4 |
5.0 residues contacted per structure ·
6 directional · 3 residues
contacted by all 4 |
mean 0.68 Å ·
peak H407 1.67 Å |
| 7 | 3 |
OCNT-2315950 — MCS highlighted · 22 atoms,
92% of the average member CCN(C(=O)c1c2ccc(Br)cc2nn1C)c1ccccc1 |
median 4.61 Å ·
0.34–4.62 Å · 3 pairs
(1 same-mode) |
2 binding modes ·
largest 2/3 |
5.3 residues contacted per structure ·
5 directional · 1 residue
contacted by all 3 |
mean 0.71 Å ·
peak H407 1.94 Å |
| 8 | 3 |
OCNT-2395537 — MCS highlighted · 19 atoms,
75% of the average member CC(=N)NC(=O)NC1(c2c(F)cccc2F)CCC1 |
median 7.53 Å ·
1.76–7.80 Å · 3 pairs
(1 same-mode) |
2 binding modes ·
largest 2/3 |
4.3 residues contacted per structure ·
6 directional · 0 residues
contacted by all 3 |
mean 0.51 Å ·
peak H407 1.73 Å |
| 9 | 3 |
OCNT-2395780 — MCS highlighted · 28 atoms,
92% of the average member CCC1CN(C(=O)c2cc(-c3cccc(C(F)(F)F)c3)nc3onc(C)c23)C1 |
median 0.76 Å ·
0.71–0.78 Å · 3 pairs
(3 same-mode) |
1 binding mode ·
largest 3/3 |
5.7 residues contacted per structure ·
13 directional · 3 residues
contacted by all 3 |
mean 0.25 Å ·
peak Y306 0.65 Å |
| 10 | 3 |
OCNT-2318048 — MCS highlighted · 15 atoms,
70% of the average member C=CNS(=O)(=O)c1cnn(C2CCC2)c1 |
median 4.74 Å ·
1.54–4.90 Å · 3 pairs
(1 same-mode) |
2 binding modes ·
largest 2/3 |
5.0 residues contacted per structure ·
3 directional · 2 residues
contacted by all 3 |
mean 0.77 Å ·
peak H407 1.60 Å |