Every compound in pxr_semi_pure_96_corrected.csv, decomposed to a shared core plus the group that actually varies, and ranked on the purity-corrected pEC50. Six congeneric series, twelve analogues each — two of them run as full 12×2 matrices.
Cross-referencing the plate against the 184 ligand-bound PXR structures in structure_ground_truth. Five plate compounds are crystallized — matched both on registration ID and, independently, on stereochemistry-stripped structure, which agree exactly. Each is marked ◆ code on its row below.
| A | Cyclohexenyl benzamides | ◆ x02813-1 1.96 Å · 2315472 | 1 further ligand on this core |
| B | Pyrazole sulfonamides | ◆ x02698-1 2.24 Å · 2318048 | — |
| C | Mandelamide piperazines | ◆ x03363-1 2.09 Å · 2308748 | 1 further ligand on this core |
| D | Piperazine ureas | no structure for any member | — |
| E | Pyrazolyl ureas | ◆ x02800-1 1.95 Å · 2318616 | 4 further ligands on this core |
| F | Sulfonylacetamide morpholines | ◆ x02793-1 2.02 Å · 2310554 | 1 further ligand on this core |
Every series except D has a bound structure of one of its own members, at 1.95–2.24 Å. Series D is the gap: 24 compounds, a quarter of the plate, and nothing crystallized on that core — its flat SAR has no structural explanation to lean on.
Coverage is richest for E, and in a complementary direction: the four extra ligands hold the phenylsulfonamide tail fixed and vary the pyrazole N-substituent — exactly the position the plate held constant while varying the tail. Between the two, that series is characterized in both directions. The extra A and C ligands are single-atom edits of the crystallized plate compound (a pyridine for the benzamide ring; a methylpyrimidine for the isopropyl), and the extra F ligand strips the 2-aryl morpholine back to a plain piperidine — a useful unsubstituted reference point.
The potency plotted throughout is the purity-corrected pEC50. It is worth knowing exactly what that correction is, because it is not an independent measurement.
Corrected potency is the observed potency shifted by the fraction of theoretical material that CAD actually found on column:
This reproduces every corrected value in the file to 10−15. It adds no new potency information — it re-expresses the same curve against the amount of compound that was really there. The uncertainty is propagated the same way: the corrected error bar is the dose–response SE and the CAD recovery SE added in quadrature.
The shift is only as good as the recovery estimate. Twelve compounds recovered under 20% of theory, and those same wells carry CAD peak-area CVs of 6–36% against a typical 2–3%. They are flagged ⚠ low recovery below.
A 1-methylcyclohex-3-enyl amide anchored to a benzenesulfonyl head group. The sulfonyl substituent is varied over nine groups; three analogues also move the sulfonyl to meta or insert a two-atom linker into the amide.
R is drawn in full, so the amide linker and the meta/para sulfonyl position are visible in the depiction — four rows share the same cyclopropylsulfonamide head and differ only there.
| R group (full acyl) | sulfonyl | ring / linker | corrected pEC50 | Emax | OCNT |
|---|---|---|---|---|---|
NH–cyclopropyl ◆ x02813-1 · 1.96 Å | para | 6.43±0.15 | 1.06 | 2315472 | |
N(Me)–cyclohexyl | para | 6.40±0.15 | 0.99 | 2469073 | |
NH–cyclopentyl | meta | 6.14±0.16 | 0.86 | 2469070 | |
NH–cyclopropyl partial | meta | 5.92±0.15 | 0.72 | 2469063 | |
NH–(1,1-dioxothiolan-3-yl) partial | para | 5.91±0.14 | 0.73 | 2469068 | |
NH–Me partial | para | 5.87±0.15 | 0.78 | 2469066 | |
pyrrolidin-1-yl | para | 5.86±0.17 | 1.25 | 2469071 | |
morpholin-4-yl | para | 5.64±0.15 | 1.09 | 2469072 | |
cyclopropyl (sulfone) partial | para | 5.45±0.15 | 0.75 | 2469067 | |
NH–cyclopropyl | para, –CH₂CH₂– | 4.88±0.18 | 0.93 | 2469065 | |
NH–cyclopropyl | para, –CH=CH– | 4.52±0.20 | 0.83 | 2469064 | |
NH–C(O)NH₂ | para | no data | 1.26 | 2469069 |
Cyclopropylsulfonamide is the best simple head group (2315472, pEC50 6.43), and a bulky N-methyl-cyclohexyl sulfonamide matches it. Potency is broadly tolerant of the sulfonyl group — nine variants span only 1.0 log — but the amide linker is not: extending the direct aryl amide to cinnamamide costs 1.9 log and to the saturated homologue 1.6 log. The sulfone (2469067) is 1.0 log worse than the matched sulfonamide, so the S–NH donor is contributing. The series lead 2315472 is the member with a bound structure (x02813-1, 1.96 Å), and none of the eleven new analogues improved on it.
A 1-cyclobutylpyrazole-4-sulfonamide on an aniline that carries a basic amine ortho to the sulfonamide NH. Only that amine is varied.
R spans the whole aniline, so the one analogue that inserts a –CH2– between ring and sulfonamide nitrogen (2469033) is visible as such rather than only noted in the linkage column.
| R group (full aniline) | ortho amine | linkage | corrected pEC50 | Emax | OCNT |
|---|---|---|---|---|---|
N(Me)ⁱPr ◆ x02698-1 · 2.24 Å | direct | 6.52±0.15 | 1.11 | 2318048 | |
N(Me)cyclopentyl partial | direct | 6.49±0.14 | 0.68 | 2469036 | |
N(Me)Ac ⚠ low recovery | direct | 6.34±0.23 | 0.83 | 2469031 | |
N(Me)Ph | direct | 6.19±0.14 | 0.94 | 2469040 | |
N(Me)ⁿPr | direct | 6.18±0.16 | 1.23 | 2469034 | |
NEt₂ | direct | 6.07±0.15 | 1.35 | 2469030 | |
N(Me)Et | direct | 6.04±0.15 | 1.12 | 2469035 | |
N(Me)ⁱPr | –CH2– | 5.81±0.15 | 0.93 | 2469033 | |
pyrrolidin-1-yl | direct | 5.78±0.15 | 1.04 | 2469039 | |
NMe₂ | direct | 5.71±0.15 | 1.32 | 2469032 | |
N(Me)SO₂Me ⚠ low recovery | direct | 5.68±0.18 | 0.87 | 2469038 | |
piperidin-1-yl | direct | 5.66±0.17 | 1.02 | 2469037 |
The tightest series on the plate — all twelve fall within 0.9 log, and every one is at or above pEC50 5.7. The trend is simple steric: going from NMe2 (5.71) up through N(Me)Et, N(Me)nPr to N(Me)iPr (6.52) and N(Me)cyclopentyl (6.49) gains ~0.8 log, while tying the amine back into a ring loses it again (pyrrolidine 5.78, piperidine 5.66). A branched acyclic amine is what this pocket wants. The crystallized lead 2318048 (x02698-1, 2.24 Å) carries exactly that N(Me)iPr group and tops the series, so the bound structure is available to explain the preference directly. Note that N(Me)Ac (6.34) rests entirely on a 3% recovery correction and should be re-tested.
A complete 12×2 matrix: twelve N-aryl piperazines, each made twice — once with an α-H mandelamide and once with the α-methyl quaternary version. This is the cleanest single-variable comparison on the plate.
R = H or CH3 at the α-carbon. Two rows replace the piperazine with a homopiperazine or a 2-methylpiperazine, noted under the aryl name.
| N-aryl | α-H | α-Me | Δ from α-Me | |
|---|---|---|---|---|
4-ⁱPr-pyrimidin-2-yl ◆ x03363-1 · 2.09 Å (α-Me) | 5.92 2469096 | 6.48 2308748 | +0.56 | |
4,6-diMe-pyrimidin-2-yl | 5.20 2469085 | 6.13 2469074 | +0.93 | |
4-Me-pyridin-2-yl partial (α-H) | 5.09 2469091 | 6.05 2469080 | +0.96 | |
4-Me-pyrimidin-2-yl partial (α-H) | 5.24 2469086 | 5.85 2469075 | +0.61 | |
4-NH₂-6-cPr-triazin-2-yl structure revised | 4.41 2469095 | 5.72 2469084 | +1.31 | |
2,3-diMe-phenyl partial (α-H)partial (α-Me) | 5.69 2469094 | 5.55 2469083 | -0.14 | |
3-Me-phenyl piperazine (+1 Me) partial (α-H) | 5.20 2469093 | 5.46 2469082 | +0.26 | |
2-ⁱPr-6-Me-pyrimidin-4-yl partial (α-Me) | 4.02 2469092 | 5.46 2469081 | +1.44 | |
3,5-diMe-phenyl | 4.35 2469089 | 5.34 2469078 | +0.99 | |
pyrimidin-2-yl homopiperazine | 5.06 2469090 | 5.33 2469079 | +0.27 | |
4-CO₂Me-pyrimidin-2-yl | 4.22 2469088 | 5.04 2469077 | +0.83 | |
pyrimidin-2-yl | 4.05 2469087 | 5.04 2469076 | +1.00 |
The α-methyl group is worth +0.75 log on average across all twelve matched pairs, and it helps in eleven of them — one of the most consistent single-atom effects in the dataset. It pays most where the aryl is otherwise poor (+1.44 on 2-iPr-6-Me-pyrimidin-4-yl, +1.31 on the aminotriazine, +1.00 on plain pyrimidine) and least where the aryl is already good, which is the signature of two partially redundant contributions to the same binding event. The single exception, 2,3-dimethylphenyl (−0.14), is the one aryl with an ortho substituent — the α-methyl and the ortho-methyl appear to compete for the same space. On the aryl side, 4-iPr-pyrimidin-2-yl is best in both halves of the matrix (6.48 / 5.92); swapping that isopropyl for a methyl ester costs ~1.4 log. That best corner of the matrix is the crystallized lead 2308748 (x03363-1, 2.09 Å) — so the α-methyl whose effect the matrix quantifies is resolved in the bound pose.
A carbonyl bridging an aryl amine or aryl ether to an N-aryl piperazine. The 24 compounds form two arms that cross at 2395416: one holds the left-hand side fixed and walks the piperazine aryl, the other does the reverse.
Arm 1 · left-hand side fixed as the 2-ethoxypyridin-3-yl urea
The left-hand side is fixed as drawn. R spans the whole diamine, so the two ring changes — a homopiperazine and a 2-methylpiperazine — are visible in the depiction.
| R group (full diamine) | piperazine aryl | corrected pEC50 | Emax | OCNT |
|---|---|---|---|---|
4-ⁱPr-pyrimidin-2-yl | 5.40±0.15 | 1.10 | 2395416 | |
2,3-diMe-phenyl partial | 5.36±0.15 | 0.67 | 2469106 | |
4,6-diMe-pyrimidin-2-yl | 5.25±0.15 | 1.03 | 2469097 | |
3-Me-phenyl piperazine (2-Me) | 5.16±0.16 | 1.11 | 2469105 | |
pyrimidin-2-yl | 5.15±0.15 | 0.98 | 2469099 | |
3,5-diMe-phenyl | 5.09±0.16 | 1.43 | 2469101 | |
pyrimidin-2-yl homopiperazine | 5.08±0.16 | 1.12 | 2469102 | |
4-NH₂-6-cPr-triazin-2-yl structure revisedpartial | 5.04±0.18 | 0.79 | 2469107 | |
4-CO₂Me-pyrimidin-2-yl ⚠ low recovery | 5.01±0.20 | 2.00 | 2469100 | |
4-Me-pyridin-2-yl | 4.97±0.18 | 1.48 | 2469103 | |
4-Me-pyrimidin-2-yl | 4.90±0.18 | 1.36 | 2469098 | |
2-ⁱPr-6-Me-pyrimidin-4-yl | 4.47±0.24 | 0.93 | 2469104 |
Remarkably flat — twelve different piperazine aryls span only 0.93 log, and eight of them sit within 0.4 log of each other. 4-iPr-pyrimidin-2-yl is again the best (5.40), and the only real loser is 2-iPr-6-Me-pyrimidin-4-yl (4.47), the same regiochemistry that underperformed in series C.
Arm 2 · right-hand side fixed as the 4-iPr-pyrimidinyl piperazine
The right-hand side is fixed as drawn; R is the left-hand side, which carries its own N or O so the bridge reads as a urea or a carbamate depending on the row.
| R group (left-hand side) | bridge | corrected pEC50 | Emax | OCNT | |
|---|---|---|---|---|---|
N(Et)Ph | urea | 6.25±0.15 | 0.94 | 2469115 | |
benzoxazin-4-yl (cyclic) | urea | 5.85±0.16 | 1.07 | 2469116 | |
O–(2-CF₃-phenyl) partial | carbamate | 5.66±0.15 | 0.65 | 2469119 | |
O–(2-MeO-phenyl) | carbamate | 5.64±0.15 | 0.95 | 2469117 | |
NH–(2-MeO-phenyl) structure revised | urea | 5.40±0.16 | 0.93 | 2469109 | |
NH–(3-ⁱPr-phenyl) ⚠ low recoverystructure revised | urea | 5.40±0.17 | 0.98 | 2469108 | |
NH–(2-EtO-pyridin-3-yl) | urea | 5.40±0.15 | 1.10 | 2395416 | |
NH–(5-ᵗBu-2-MeO-phenyl) structure revised | urea | 5.15±0.20 | 0.95 | 2469112 | |
NH–(3-EtO-phenyl) structure revised | urea | 5.13±0.22 | 0.89 | 2469111 | |
O–(2,4-diF-phenyl) ⚠ low recovery | carbamate | 5.09±0.26 | 1.08 | 2469118 | |
NH–(benzodioxol-4-yl) ⚠ low recoverystructure revised | urea | 4.99±0.22 | 1.12 | 2469114 | |
NH–(pyridin-3-yl) structure revised | urea | 4.88±0.20 | 0.88 | 2469113 | |
NH–(6-Me-pyridin-3-yl) ⚠ low recoverystructure revised | urea | 4.82±0.23 | 1.26 | 2469110 |
The left-hand side matters more, spanning 1.4 log. Two things read clearly: a tertiary urea beats every secondary one (N-ethyl-N-phenyl at 6.25 is the best compound in the series, and the cyclic benzoxazine urea is second at 5.85, both on clean >100% recoveries), and an ortho substituent on the aniline is worth roughly half a log (2-MeO-phenyl 5.40 vs. pyridin-3-yl 4.88). Carbamates are competitive with the secondary ureas but not with the tertiary ones. Four of the weakest rows carry large purity corrections, so the bottom of this ranking is the softest data on the plate. D is also the one series with no crystal structure — not for its lead 2395416, and not for any other compound on this core — so unlike every other series here, its flat SAR cannot be checked against a bound pose.
A 1-(4-fluoro-2-methylphenyl)pyrazol-3-yl urea — constant across all twelve — with a sulfonamide-bearing tail. The tail is the only variable.
The Emax column is the story in this series: every member is a partial agonist.
| tail | corrected pEC50 | Emax | OCNT | |
|---|---|---|---|---|
NH(CH₂)₂SO₂NH–Ph ◆ x02800-1 · 1.95 Åpartial | 6.68±0.15 | 0.63 | 2318616 | |
NH(CH₂)₂SO₂NH–(pyridin-2-yl) ⚠ low recoverypartial | 6.59±0.22 | 0.70 | 2469042 | |
NH(CH₂)₂SO₂NH–Bn partial | 6.25±0.14 | 0.63 | 2469043 | |
NH(CH₂)₂SO₂–indolin-1-yl partial | 6.12±0.14 | 0.63 | 2469047 | |
NH(CH₂)₂NH–SO₂Ph (reversed) partial | 6.09±0.15 | 0.64 | 2469041 | |
NH(CH₂)₂C(O)NH–Ph partial | 5.86±0.17 | 0.68 | 2469050 | |
3-(PhNHSO₂)azetidin-1-yl partial | 5.86±0.17 | 0.65 | 2469045 | |
NH(CH₂)₂SO₂NH–(2-Br-pyridin-4-yl) ⚠ low recoverypartial | 5.84±0.15 | 0.66 | 2469044 | |
NH(CH₂)₂SO₂NH–(3-F-4-Me-phenyl) partial | 5.70±0.19 | 0.66 | 2469049 | |
NHCH(Me)CH₂C(O)NH–Ph | 5.69±0.26 | 0.80 | 2469051 | |
4-(PhSO₂)piperazin-1-yl | 5.49±0.25 | 0.83 | 2469046 | |
NH(CH₂)₂S(=NH)(O)Ph (sulfoximine) ⚠ low recovery | no data | — | 2469048 |
This series is pharmacologically distinct from the other five. Its Emax values cluster at 0.63–0.83 (mean 0.68) while every other series averages 0.94–1.07 — these are partial agonists, and the effect tracks the scaffold rather than the tail, so it is the constant pyrazolyl urea head that caps efficacy. If the goal is a ceiling on PXR activation rather than raw potency, this is the series to work on. Within the tail, the two-carbon sulfonamide is the preferred geometry: a plain phenyl sulfonamide leads (6.68), N-benzyl and indoline are close behind, and both conformationally restrained versions — the azetidine (5.86) and the piperazine (5.49) — give the tail up without gaining anything. Replacing the sulfonamide with a carboxamide costs 0.8 log. This is the best-characterized series structurally: the lead 2318616 is crystallized (x02800-1, 1.95 Å) and four further structures keep that phenylsulfonamide tail while varying the pyrazole N-substituent — the one position the plate held fixed.
A 3,4-dimethylphenylsulfonyl acetamide on a 2-aryl morpholine. The morpholine and its aryl vary together.
Most rows are 2-aryl morpholines; three change the ring itself — a 1,4-oxazepane homologue, a 5,5-dimethyl regioisomer and a spiro-indane.
| amine | corrected pEC50 | Emax | OCNT | |
|---|---|---|---|---|
2-(4-CF₃-phenyl)morpholine ⚠ low recoverypartial | 6.97±0.15 | 0.79 | 2469052 | |
5,5-diMe-2-Ph-morpholine | 6.89±0.15 | 1.18 | 2469060 | |
2-(3-Br-4-F-phenyl)morpholine | 6.72±0.15 | 1.06 | 2469054 | |
2-(4-OCHF₂-phenyl)morpholine | 6.64±0.14 | 0.93 | 2469057 | |
2-(4-F-phenyl)-1,4-oxazepane | 6.62±0.14 | 1.02 | 2469056 | |
2-(4-Me-phenyl)morpholine | 6.62±0.14 | 1.05 | 2469053 | |
2-(4-F-phenyl)morpholine ◆ x02793-1 · 2.02 Å | 6.32±0.15 | 1.06 | 2310554 | |
2-(3-F-phenyl)morpholine | 6.26±0.15 | 1.13 | 2469055 | |
2-(2-F-phenyl)morpholine | 5.78±0.14 | 1.18 | 2469059 | |
spiro[indane-morpholine] ⚠ low recovery | 5.56±0.16 | 0.80 | 2469062 | |
2-(pyridin-2-yl)morpholine | 5.50±0.15 | 0.87 | 2469058 | |
2-Me-2-Ph-morpholine | 5.29±0.16 | 1.03 | 2469061 |
The most potent series on the plate — median pEC50 6.47, and it holds the top compound overall. The SAR is a clean read on the 2-aryl group: para substitution is strongly preferred and the effect tracks lipophilicity (4-CF3 6.97 > 4-OCHF2 6.64 ≈ 4-Me 6.62 > 4-F 6.32), while walking that same fluorine around the ring costs potency at every step (para 6.32 → meta 6.26 → ortho 5.78). Two changes are clearly disfavoured: swapping the phenyl for 2-pyridyl loses 0.8 log, and methylating the morpholine 2-position to a quaternary centre loses 1.3 log — the opposite of what the same substitution did in series C. The 5,5-dimethyl regioisomer (6.89) and the oxazepane homologue (6.62) are both well tolerated, so the ring itself has room to move. F is the series that most clearly moved: its crystallized lead 2310554 (x02793-1, 2.02 Å) sits only 7th of twelve, beaten by six analogues and by 0.64 log at the top — and a bound pose of the parent 4-fluorophenyl exists to rationalize why.
Six compounds change rank materially on the purity correction alone: 2469031, 2469042 and 2469118 move 1.35–1.51 log on recoveries of 3–5%, and 2469052, 2469110 and 2469114 move 0.75–0.92 log. Three of those — 2469031 in B, 2469042 in E, 2469052 in F — are currently top-3 in their series. Confirming them on purified material is the highest-value follow-up on the plate, because three of the six SAR conclusions above lean on compounds whose potency is known only through a 20–30× extrapolation.
Built from pxr_semi_pure_96_corrected.csv (96 rows). Series assigned by substructure match; R-groups obtained by cutting each molecule at the bonds that define its series, then verified by reconstruction. Potency is the corrected pEC50 with its combined dose–response and recovery error; Emax is the normalized value as supplied. Structure depictions are drawn from the corrected SMILES.
For each of the five crystallized leads, every analogue in that series was given a conformer ensemble (ETKDG, MMFF-minimised), superimposed on the crystal ligand across its shared scaffold, and the single conformer with the best combined shape + feature overlap was kept. Series A, B and C used 200 conformers; the two flexible series, E and F, used 1000. What you see below is that best pose for each analogue, inside the protein it was aligned into.
Conformers were superimposed on the shared scaffold (the maximum common substructure with the crystal ligand, 18–24 heavy atoms of a ~25-atom molecule) rather than by free shape fitting. A free shape fit maximises overlap without regard to chemistry, and on these elongated molecules it happily flips an analogue end-to-end: it raised the mean core RMSD in series F from 0.8–1.8 Å to 3.3 Å, superimposing tails onto heads. Anchoring the superposition and still selecting the conformer on shape + feature overlap keeps the stated criterion and gives poses a chemist would recognise.
Two of the five deposited structures — series C and E — carry two copies of the ligand, and the copy nominated in chains_df.csv is badly modelled in both: a carbonyl oxygen 1.30 Å from Arg410 in C, a urea nitrogen 1.43 Å from Met323 in E. The second copy is clean in both cases, and that is the one used here.
Each analogue carries its core RMSD to the crystal scaffold and its shape+feature score, shown in the ligand list and summarised by a dot: under 1.2 Å, 1.2–2.0 Å, over 2.0 Å. Series A, B, C and F place every analogue under 2 Å; series E places 10 of 11.
Going to 1000 conformers rescued F but not E. Both carry 6–7 rotatable bonds and both were clearly under-sampled at 200. Series F responded: mean overlap 1.03 → 1.21, mean core RMSD 1.80 → 1.39 Å, and analogues under 2 Å went from 7 of 11 to all 11 — its four worst poses all improved by 0.2–0.4 overlap units. Series E barely moved (0.75 → 0.80; core RMSD 1.56 → 1.55 Å), so its low overlap is not a sampling problem — these tails simply cannot reproduce the extended crystal conformation while their ureas stay superimposed. Read E’s poses as scaffold-anchored only, and its tails as unresolved.
One case moved the wrong way: 2469049 gained overlap (0.76 → 0.85) but lost scaffold agreement (1.90 → 2.98 Å) — the larger ensemble contained a conformer that overlaps the crystal ligand better while sitting worse on the shared core. That is inherent to selecting on overlap, and it is the one red dot in series E.
Hydrogen bonds are heavy-atom N/O…N/O contacts within 3.5 Å — these structures carry no hydrogens, so donor geometry is not checked. For the crystal ligands the contacts land on PXR’s canonical polar residues (Ser247, Gln285, His327, His407, Arg410), which is a good sign the geometry is right. On an aligned analogue an H-bond is a prediction, not an observation.