The pregnane X receptor is the body’s sensor for foreign chemicals, and a master regulator of their disposal: activate it and it transcribes the machinery that clears them, CYP3A4 above all, along with other P450s, conjugating enzymes, and efflux transporters such as P-glycoprotein. A compound that turns PXR on therefore speeds up its own clearance, and that of anything taken alongside it, which means exposure lost before the drug reaches its target and a well-known route to serious drug–drug interactions.
That makes PXR a textbook member of the Avoid-ome, the enzymes, transporters and receptors a discovery programme would rather its compounds left alone. The 96 compounds here come from a high-throughput PXR campaign: made at microscale, purified only partially, assayed as made. This report reads the SAR out of that plate as six congeneric series, each an expansion around a single lead, decomposed to a shared core plus the group that varies. For the five leads crystallized with PXR, every analogue is overlaid on that experimental pose. Potency here is a liability, not a goal.
The data and the methods behind it are described in “Navigating PXR chemical space with high-throughput chemistry and standard-free quantification”.
Where these 96 compounds came from, and why the potency column has the shape it does.
PXR is a xenobiotic sensor, and in drug discovery it is an antitarget: a receptor you want your compound to leave alone. It is unusually promiscuous, and the public record is thin: on the order of 800 trustworthy EC50 values in ChEMBL, drawn from ~148 papers under conditions that were never standardised. Octant and OpenADMET set out to fix that by measuring a large, internally consistent slice of PXR chemical space in one assay.
The chemistry was high-throughput: reference actives were cut back to carboxylic-acid cores and recoupled against a 1,536-member amine fragment set, then read out in an in-house cell-based PXR agonism reporter assay over nine-point dose–response. This file is the 96-compound “semi-pure” plate from that campaign: microscale synthesis, partial purification, straight into the assay. It resolves into six congeneric series, each an expansion around one previously registered lead.
Read the direction of merit backwards. Everything below ranks compounds by potency because that is what was measured, but on an antitarget a high pEC50 is the liability.
And nothing here is clean. This is an expansion around known actives, so every compound on the plate activates PXR, at a median of 2.2 µM, with only 13 of 94 weaker than 10 µM. Series C and D are the least potent, but that means medians of 5.2 and 7.3 µM, which is still a real induction liability rather than a safe place to be. The plate is useful for direction, meaning which structural changes lower PXR activation, not for picking a winner.
Semi-pure material means the amount of compound in the well is not the amount you weighed out. Rather than build a calibration standard for every analogue, the method uses charged aerosol detection, whose response is close to structure-independent, with one universal conversion:
That constant holds exactly across all 96 rows here. Yield follows as the ratio to theoretical mass, and the potency is rescaled by it. That is where the corrected pEC50 comes from, and why it is a re-expression of the same curve rather than a second measurement.
The price is a floor on every error bar. Individual compounds scatter around the universal slope by 31.7%, and that term is applied identically to all 96, so no corrected value can be more precise than ±0.138 log. Median corrected SE on this plate is 0.154, of which ~80% of the variance is that one constant. Treat gaps under ~0.2 log as noise, however tidy the ranking looks.
Cross-referencing the plate against the 184 ligand-bound PXR structures in structure_ground_truth. Five plate compounds are crystallized, matched both on registration ID and, independently, on stereochemistry-stripped structure, which agree exactly. Each is marked ◆ code on its row below.
| A | Cyclohexenyl benzamides | ◆ x02813-1 1.96 Å · 2315472 | 1 further ligand on this core |
| B | Pyrazole sulfonamides | ◆ x02698-1 2.24 Å · 2318048 | – |
| C | Mandelamide piperazines | ◆ x03363-1 2.09 Å · 2308748 | 1 further ligand on this core |
| D | Piperazine ureas | no structure for any member | – |
| E | Pyrazolyl ureas | ◆ x02800-1 1.95 Å · 2318616 | 4 further ligands on this core |
| F | Sulfonylacetamide morpholines | ◆ x02793-1 2.02 Å · 2310554 | 1 further ligand on this core |
Every series except D has a bound structure of one of its own members, at 1.95–2.24 Å. Series D is the gap: 24 compounds, a quarter of the plate, and nothing crystallized on that core, so its flat SAR has no structural explanation to lean on.
Coverage is richest for E, and in a complementary direction: the four extra ligands hold the phenylsulfonamide tail fixed and vary the pyrazole N-substituent, exactly the position the plate held constant while varying the tail. Between the two, that series is characterized in both directions. The extra A and C ligands are single-atom edits of the crystallized plate compound (a pyridine for the benzamide ring; a methylpyrimidine for the isopropyl), and the extra F ligand strips the 2-aryl morpholine back to a plain piperidine, a useful unsubstituted reference point.
The potency plotted throughout is the purity-corrected pEC50. It is worth knowing exactly what that correction is, because it is not an independent measurement.
Corrected potency is the observed potency shifted by the fraction of theoretical material that CAD actually found on column:
This reproduces every corrected value in the file to 10−15. It adds no new potency information, since it re-expresses the same curve against the amount of compound that was really there. The uncertainty is propagated the same way. The corrected error bar is the dose–response fit SE and the CAD yield SE in quadrature, and that yield term is itself the peak-area CV and the 31.74% inter-analyte slope CV in quadrature. Both reproduce to 10−16. Because the slope term is a constant, it sets the ±0.138 log floor noted above.
The shift is only as good as the recovery estimate. Twelve compounds recovered under 20% of theory, and those same wells carry CAD peak-area CVs of 6–36% against a typical 2–3%. They are flagged ⚠ low recovery below.
A 1-methylcyclohex-3-enyl amide anchored to a benzenesulfonyl head group. The sulfonyl substituent is varied over nine groups; three analogues also move the sulfonyl to meta or insert a two-atom linker into the amide.
R is drawn in full, so the amide linker and the meta/para sulfonyl position are visible in the depiction, since four rows share the same cyclopropylsulfonamide head and differ only there.
| R group (full acyl) | sulfonyl | ring / linker | corrected pEC50 | Emax | OCNT |
|---|---|---|---|---|---|
NH–cyclopropyl ◆ x02813-1 · 1.96 Å | para | 6.43±0.15 | 1.06 | 2315472 | |
N(Me)–cyclohexyl | para | 6.40±0.15 | 0.99 | 2469073 | |
NH–cyclopentyl | meta | 6.14±0.16 | 0.86 | 2469070 | |
NH–cyclopropyl partial | meta | 5.92±0.15 | 0.72 | 2469063 | |
NH–(1,1-dioxothiolan-3-yl) partial | para | 5.91±0.14 | 0.73 | 2469068 | |
NH–Me partial | para | 5.87±0.15 | 0.78 | 2469066 | |
pyrrolidin-1-yl | para | 5.86±0.17 | 1.25 | 2469071 | |
morpholin-4-yl | para | 5.64±0.15 | 1.09 | 2469072 | |
cyclopropyl (sulfone) partial | para | 5.45±0.15 | 0.75 | 2469067 | |
NH–cyclopropyl | para, –CH₂CH₂– | 4.88±0.18 | 0.93 | 2469065 | |
NH–cyclopropyl | para, –CH=CH– | 4.52±0.20 | 0.83 | 2469064 | |
NH–C(O)NH₂ | para | no data | 1.26 | 2469069 |
Cyclopropylsulfonamide is the best simple head group (2315472, pEC50 6.43), and a bulky N-methyl-cyclohexyl sulfonamide matches it. Potency is broadly tolerant of the sulfonyl group, with nine variants spanning only 1.0 log, but the amide linker is not: extending the direct aryl amide to cinnamamide costs 1.9 log and to the saturated homologue 1.6 log. The sulfone (2469067) is 1.0 log worse than the matched sulfonamide, so the S–NH donor is contributing. The series lead 2315472 is the member with a bound structure (x02813-1, 1.96 Å), and none of the eleven new analogues improved on it.
A 1-cyclobutylpyrazole-4-sulfonamide on an aniline that carries a basic amine ortho to the sulfonamide NH. Only that amine is varied.
R spans the whole aniline, so the one analogue that inserts a –CH2– between ring and sulfonamide nitrogen (2469033) is visible as such rather than only noted in the linkage column.
| R group (full aniline) | ortho amine | linkage | corrected pEC50 | Emax | OCNT |
|---|---|---|---|---|---|
N(Me)ⁱPr ◆ x02698-1 · 2.24 Å | direct | 6.52±0.15 | 1.11 | 2318048 | |
N(Me)cyclopentyl partial | direct | 6.49±0.14 | 0.68 | 2469036 | |
N(Me)Ac ⚠ low recovery | direct | 6.34±0.23 | 0.83 | 2469031 | |
N(Me)Ph | direct | 6.19±0.14 | 0.94 | 2469040 | |
N(Me)ⁿPr | direct | 6.18±0.16 | 1.23 | 2469034 | |
NEt₂ | direct | 6.07±0.15 | 1.35 | 2469030 | |
N(Me)Et | direct | 6.04±0.15 | 1.12 | 2469035 | |
N(Me)ⁱPr | –CH2– | 5.81±0.15 | 0.93 | 2469033 | |
pyrrolidin-1-yl | direct | 5.78±0.15 | 1.04 | 2469039 | |
NMe₂ | direct | 5.71±0.15 | 1.32 | 2469032 | |
N(Me)SO₂Me ⚠ low recovery | direct | 5.68±0.18 | 0.87 | 2469038 | |
piperidin-1-yl | direct | 5.66±0.17 | 1.02 | 2469037 |
The tightest series on the plate: all twelve fall within 0.9 log, and every one is at or above pEC50 5.7. The trend is simple steric: going from NMe2 (5.71) up through N(Me)Et, N(Me)nPr to N(Me)iPr (6.52) and N(Me)cyclopentyl (6.49) gains ~0.8 log, while tying the amine back into a ring loses it again (pyrrolidine 5.78, piperidine 5.66). A branched acyclic amine is what this pocket wants. The crystallized lead 2318048 (x02698-1, 2.24 Å) carries exactly that N(Me)iPr group and tops the series, so the bound structure is available to explain the preference directly. Note that N(Me)Ac (6.34) rests entirely on a 3% recovery correction and should be re-tested.
A complete 12×2 matrix: twelve N-aryl piperazines, each made twice, once with an α-H mandelamide and once with the α-methyl quaternary version. This is the cleanest single-variable comparison on the plate.
R = H or CH3 at the α-carbon. Two rows replace the piperazine with a homopiperazine or a 2-methylpiperazine, noted under the aryl name.
| N-aryl | α-H | α-Me | Δ from α-Me | |
|---|---|---|---|---|
4-ⁱPr-pyrimidin-2-yl ◆ x03363-1 · 2.09 Å (α-Me) | 5.92 2469096 | 6.48 2308748 | +0.56 | |
4,6-diMe-pyrimidin-2-yl | 5.20 2469085 | 6.13 2469074 | +0.93 | |
4-Me-pyridin-2-yl partial (α-H) | 5.09 2469091 | 6.05 2469080 | +0.96 | |
4-Me-pyrimidin-2-yl partial (α-H) | 5.24 2469086 | 5.85 2469075 | +0.61 | |
4-NH₂-6-cPr-triazin-2-yl structure revised | 4.41 2469095 | 5.72 2469084 | +1.31 | |
2,3-diMe-phenyl partial (α-H)partial (α-Me) | 5.69 2469094 | 5.55 2469083 | -0.14 | |
3-Me-phenyl piperazine (+1 Me) partial (α-H) | 5.20 2469093 | 5.46 2469082 | +0.26 | |
2-ⁱPr-6-Me-pyrimidin-4-yl partial (α-Me) | 4.02 2469092 | 5.46 2469081 | +1.44 | |
3,5-diMe-phenyl | 4.35 2469089 | 5.34 2469078 | +0.99 | |
pyrimidin-2-yl homopiperazine | 5.06 2469090 | 5.33 2469079 | +0.27 | |
4-CO₂Me-pyrimidin-2-yl | 4.22 2469088 | 5.04 2469077 | +0.83 | |
pyrimidin-2-yl | 4.05 2469087 | 5.04 2469076 | +1.00 |
The α-methyl group is worth +0.75 log on average across all twelve matched pairs, and it helps in eleven of them, one of the most consistent single-atom effects in the dataset. It pays most where the aryl is otherwise poor (+1.44 on 2-iPr-6-Me-pyrimidin-4-yl, +1.31 on the aminotriazine, +1.00 on plain pyrimidine) and least where the aryl is already good, which is the signature of two partially redundant contributions to the same binding event. The single exception, 2,3-dimethylphenyl (−0.14), is the one aryl with an ortho substituent, and the α-methyl and the ortho-methyl appear to compete for the same space. On the aryl side, 4-iPr-pyrimidin-2-yl is best in both halves of the matrix (6.48 / 5.92); swapping that isopropyl for a methyl ester costs ~1.4 log. That best corner of the matrix is the crystallized lead 2308748 (x03363-1, 2.09 Å), so the α-methyl whose effect the matrix quantifies is resolved in the bound pose.
A carbonyl bridging an aryl amine or aryl ether to an N-aryl piperazine. The 24 compounds form two arms that cross at 2395416: one holds the left-hand side fixed and walks the piperazine aryl, the other does the reverse.
Arm 1 · left-hand side fixed as the 2-ethoxypyridin-3-yl urea
The left-hand side is fixed as drawn. R spans the whole diamine, so the two ring changes, a homopiperazine and a 2-methylpiperazine, are visible in the depiction.
| R group (full diamine) | piperazine aryl | corrected pEC50 | Emax | OCNT |
|---|---|---|---|---|
4-ⁱPr-pyrimidin-2-yl | 5.40±0.15 | 1.10 | 2395416 | |
2,3-diMe-phenyl partial | 5.36±0.15 | 0.67 | 2469106 | |
4,6-diMe-pyrimidin-2-yl | 5.25±0.15 | 1.03 | 2469097 | |
3-Me-phenyl piperazine (2-Me) | 5.16±0.16 | 1.11 | 2469105 | |
pyrimidin-2-yl | 5.15±0.15 | 0.98 | 2469099 | |
3,5-diMe-phenyl | 5.09±0.16 | 1.43 | 2469101 | |
pyrimidin-2-yl homopiperazine | 5.08±0.16 | 1.12 | 2469102 | |
4-NH₂-6-cPr-triazin-2-yl structure revisedpartial | 5.04±0.18 | 0.79 | 2469107 | |
4-CO₂Me-pyrimidin-2-yl ⚠ low recovery | 5.01±0.20 | 2.00 | 2469100 | |
4-Me-pyridin-2-yl | 4.97±0.18 | 1.48 | 2469103 | |
4-Me-pyrimidin-2-yl | 4.90±0.18 | 1.36 | 2469098 | |
2-ⁱPr-6-Me-pyrimidin-4-yl | 4.47±0.24 | 0.93 | 2469104 |
Remarkably flat: twelve different piperazine aryls span only 0.93 log, and eight of them sit within 0.4 log of each other. 4-iPr-pyrimidin-2-yl is again the best (5.40), and the only real loser is 2-iPr-6-Me-pyrimidin-4-yl (4.47), the same regiochemistry that underperformed in series C.
Arm 2 · right-hand side fixed as the 4-iPr-pyrimidinyl piperazine
The right-hand side is fixed as drawn; R is the left-hand side, which carries its own N or O so the bridge reads as a urea or a carbamate depending on the row.
| R group (left-hand side) | bridge | corrected pEC50 | Emax | OCNT | |
|---|---|---|---|---|---|
N(Et)Ph | urea | 6.25±0.15 | 0.94 | 2469115 | |
benzoxazin-4-yl (cyclic) | urea | 5.85±0.16 | 1.07 | 2469116 | |
O–(2-CF₃-phenyl) partial | carbamate | 5.66±0.15 | 0.65 | 2469119 | |
O–(2-MeO-phenyl) | carbamate | 5.64±0.15 | 0.95 | 2469117 | |
NH–(2-MeO-phenyl) structure revised | urea | 5.40±0.16 | 0.93 | 2469109 | |
NH–(3-ⁱPr-phenyl) ⚠ low recoverystructure revised | urea | 5.40±0.17 | 0.98 | 2469108 | |
NH–(2-EtO-pyridin-3-yl) | urea | 5.40±0.15 | 1.10 | 2395416 | |
NH–(5-ᵗBu-2-MeO-phenyl) structure revised | urea | 5.15±0.20 | 0.95 | 2469112 | |
NH–(3-EtO-phenyl) structure revised | urea | 5.13±0.22 | 0.89 | 2469111 | |
O–(2,4-diF-phenyl) ⚠ low recovery | carbamate | 5.09±0.26 | 1.08 | 2469118 | |
NH–(benzodioxol-4-yl) ⚠ low recoverystructure revised | urea | 4.99±0.22 | 1.12 | 2469114 | |
NH–(pyridin-3-yl) structure revised | urea | 4.88±0.20 | 0.88 | 2469113 | |
NH–(6-Me-pyridin-3-yl) ⚠ low recoverystructure revised | urea | 4.82±0.23 | 1.26 | 2469110 |
The left-hand side matters more, spanning 1.4 log. Two things read clearly: a tertiary urea beats every secondary one (N-ethyl-N-phenyl at 6.25 is the best compound in the series, and the cyclic benzoxazine urea is second at 5.85, both on clean >100% recoveries), and an ortho substituent on the aniline is worth roughly half a log (2-MeO-phenyl 5.40 vs. pyridin-3-yl 4.88). Carbamates are competitive with the secondary ureas but not with the tertiary ones. Four of the weakest rows carry large purity corrections, so the bottom of this ranking is the softest data on the plate. D is also the one series with no crystal structure, not for its lead 2395416 and not for any other compound on this core, so unlike every other series here, its flat SAR cannot be checked against a bound pose.
A 1-(4-fluoro-2-methylphenyl)pyrazol-3-yl urea, constant across all twelve, with a sulfonamide-bearing tail. The tail is the only variable.
The Emax column is the story in this series: every member is a partial agonist.
| tail | corrected pEC50 | Emax | OCNT | |
|---|---|---|---|---|
NH(CH₂)₂SO₂NH–Ph ◆ x02800-1 · 1.95 Åpartial | 6.68±0.15 | 0.63 | 2318616 | |
NH(CH₂)₂SO₂NH–(pyridin-2-yl) ⚠ low recoverypartial | 6.59±0.22 | 0.70 | 2469042 | |
NH(CH₂)₂SO₂NH–Bn partial | 6.25±0.14 | 0.63 | 2469043 | |
NH(CH₂)₂SO₂–indolin-1-yl partial | 6.12±0.14 | 0.63 | 2469047 | |
NH(CH₂)₂NH–SO₂Ph (reversed) partial | 6.09±0.15 | 0.64 | 2469041 | |
NH(CH₂)₂C(O)NH–Ph partial | 5.86±0.17 | 0.68 | 2469050 | |
3-(PhNHSO₂)azetidin-1-yl partial | 5.86±0.17 | 0.65 | 2469045 | |
NH(CH₂)₂SO₂NH–(2-Br-pyridin-4-yl) ⚠ low recoverypartial | 5.84±0.15 | 0.66 | 2469044 | |
NH(CH₂)₂SO₂NH–(3-F-4-Me-phenyl) partial | 5.70±0.19 | 0.66 | 2469049 | |
NHCH(Me)CH₂C(O)NH–Ph | 5.69±0.26 | 0.80 | 2469051 | |
4-(PhSO₂)piperazin-1-yl | 5.49±0.25 | 0.83 | 2469046 | |
NH(CH₂)₂S(=NH)(O)Ph (sulfoximine) ⚠ low recovery | no data | – | 2469048 |
This series is pharmacologically distinct from the other five. Its Emax values cluster at 0.63–0.83 (mean 0.68) while every other series averages 0.94–1.07. These are partial agonists, and the effect tracks the scaffold rather than the tail, so it is the constant pyrazolyl urea head that caps efficacy. If the goal is a ceiling on PXR activation rather than raw potency, this is the series to work on. Within the tail, the two-carbon sulfonamide is the preferred geometry: a plain phenyl sulfonamide leads (6.68), N-benzyl and indoline are close behind, and both conformationally restrained versions, the azetidine (5.86) and the piperazine (5.49), give the tail up without gaining anything. Replacing the sulfonamide with a carboxamide costs 0.8 log. This is the best-characterized series structurally: the lead 2318616 is crystallized (x02800-1, 1.95 Å) and four further structures keep that phenylsulfonamide tail while varying the pyrazole N-substituent, the one position the plate held fixed.
A 3,4-dimethylphenylsulfonyl acetamide on a 2-aryl morpholine. The morpholine and its aryl vary together.
Most rows are 2-aryl morpholines; three change the ring itself: a 1,4-oxazepane homologue, a 5,5-dimethyl regioisomer and a spiro-indane.
| amine | corrected pEC50 | Emax | OCNT | |
|---|---|---|---|---|
2-(4-CF₃-phenyl)morpholine ⚠ low recoverypartial | 6.97±0.15 | 0.79 | 2469052 | |
5,5-diMe-2-Ph-morpholine | 6.89±0.15 | 1.18 | 2469060 | |
2-(3-Br-4-F-phenyl)morpholine | 6.72±0.15 | 1.06 | 2469054 | |
2-(4-OCHF₂-phenyl)morpholine | 6.64±0.14 | 0.93 | 2469057 | |
2-(4-F-phenyl)-1,4-oxazepane | 6.62±0.14 | 1.02 | 2469056 | |
2-(4-Me-phenyl)morpholine | 6.62±0.14 | 1.05 | 2469053 | |
2-(4-F-phenyl)morpholine ◆ x02793-1 · 2.02 Å | 6.32±0.15 | 1.06 | 2310554 | |
2-(3-F-phenyl)morpholine | 6.26±0.15 | 1.13 | 2469055 | |
2-(2-F-phenyl)morpholine | 5.78±0.14 | 1.18 | 2469059 | |
spiro[indane-morpholine] ⚠ low recovery | 5.56±0.16 | 0.80 | 2469062 | |
2-(pyridin-2-yl)morpholine | 5.50±0.15 | 0.87 | 2469058 | |
2-Me-2-Ph-morpholine | 5.29±0.16 | 1.03 | 2469061 |
The most potent series on the plate, at a median pEC50 of 6.47, and it holds the top compound overall. The SAR is a clean read on the 2-aryl group: para substitution is strongly preferred and the effect tracks lipophilicity (4-CF3 6.97 > 4-OCHF2 6.64 ≈ 4-Me 6.62 > 4-F 6.32), while walking that same fluorine around the ring costs potency at every step (para 6.32 → meta 6.26 → ortho 5.78). Two changes are clearly disfavoured: swapping the phenyl for 2-pyridyl loses 0.8 log, and methylating the morpholine 2-position to a quaternary centre loses 1.3 log, the opposite of what the same substitution did in series C. The 5,5-dimethyl regioisomer (6.89) and the oxazepane homologue (6.62) are both well tolerated, so the ring itself has room to move. F is the series that most clearly moved: its crystallized lead 2310554 (x02793-1, 2.02 Å) sits only 7th of twelve, beaten by six analogues and by 0.64 log at the top, and a bound pose of the parent 4-fluorophenyl exists to rationalize why.
Six compounds change rank materially on the purity correction alone: 2469031, 2469042 and 2469118 move 1.35–1.51 log on recoveries of 3–5%, and 2469052, 2469110 and 2469114 move 0.75–0.92 log. Three of those (2469031 in B, 2469042 in E, 2469052 in F) are currently top-3 in their series. Confirming them on purified material is the highest-value follow-up on the plate, because three of the six SAR conclusions above lean on compounds whose potency is known only through a 20–30× extrapolation.
Built from pxr_semi_pure_96_corrected.csv (96 rows). Series assigned by substructure match; R-groups obtained by cutting each molecule at the bonds that define its series, then verified by reconstruction. Potency is the corrected pEC50 with its combined dose–response and recovery error; Emax is the normalized value as supplied. Structure depictions are drawn from the corrected SMILES.
For each of the five crystallized leads, every analogue in that series was given a conformer ensemble (ETKDG, MMFF-minimised), superimposed on the crystal ligand across its shared scaffold, and the single conformer with the best combined shape + feature overlap was kept. Series A, B and C used 200 conformers; the two flexible series, E and F, used 1000. What you see below is that best pose for each analogue, inside the protein it was aligned into.
Conformers were superimposed on the shared scaffold (the maximum common substructure with the crystal ligand, 18–24 heavy atoms of a ~25-atom molecule) rather than by free shape fitting. A free shape fit maximises overlap without regard to chemistry, and on these elongated molecules it happily flips an analogue end-to-end: it raised the mean core RMSD in series F from 0.8–1.8 Å to 3.3 Å, superimposing tails onto heads. Anchoring the superposition and still selecting the conformer on shape + feature overlap keeps the stated criterion and gives poses a chemist would recognise.
Two of the five deposited structures, series C and E, carry two copies of the ligand, and the copy nominated in chains_df.csv is badly modelled in both: a carbonyl oxygen 1.30 Å from Arg410 in C, a urea nitrogen 1.43 Å from Met323 in E. The second copy is clean in both cases, and that is the one used here.
Each analogue carries its core RMSD to the crystal scaffold and its shape+feature score, shown in the ligand list and summarised by a dot: under 1.2 Å, 1.2–2.0 Å, over 2.0 Å. Series A, B, C and F place every analogue under 2 Å; series E places 10 of 11.
Going to 1000 conformers rescued F but not E. Both carry 6–7 rotatable bonds and both were clearly under-sampled at 200. Series F responded: mean overlap 1.03 → 1.21, mean core RMSD 1.80 → 1.39 Å, and analogues under 2 Å went from 7 of 11 to all 11, with its four worst poses all improving by 0.2–0.4 overlap units. Series E barely moved (0.75 → 0.80; core RMSD 1.56 → 1.55 Å), so its low overlap is not a sampling problem. These tails simply cannot reproduce the extended crystal conformation while their ureas stay superimposed. Read E’s poses as scaffold-anchored only, and its tails as unresolved.
One case moved the wrong way: 2469049 gained overlap (0.76 → 0.85) but lost scaffold agreement (1.90 → 2.98 Å), because the larger ensemble contained a conformer that overlaps the crystal ligand better while sitting worse on the shared core. That is inherent to selecting on overlap, and it is the one red dot in series E.
Hydrogen bonds are heavy-atom N/O…N/O contacts within 3.5 Å. These structures carry no hydrogens, so donor geometry is not checked. For the crystal ligands the contacts land on PXR’s canonical polar residues (Ser247, Gln285, His327, His407, Arg410), which is a good sign the geometry is right. On an aligned analogue an H-bond is a prediction, not an observation.
We would like to thank our funders for their support of OpenADMET, in particular ARPA-H, Radial (part of the Astera Institute), Schrödinger Inc, and the Gates Foundation. We would also like to thank our partners Enamine, HuggingFace, OpenEye, CDD Vault, Discovery Life Sciences, and the beamline staff at NSLS-II for their support.
This work is supported by the Advanced Research Projects Agency for Health (ARPA-H) under AVOID-OME, and Award Number 1AY1AX000035. The contents are those of the authors. They may not reflect the policies of the Department of Health and Human Services or the U.S. government. The content is solely the responsibility of the authors and does not necessarily represent the official views of the Advanced Research Projects Agency for Health.