The Cohort B results of the WILLOW trial have prompted a methodological debate over how lupus trials distinguish between a drug that fails to alter disease biology and an endpoint that is not structured to detect such an alteration. A correspondence published in The Lancet argues that these two questions are often conflated in lupus research, and that the WILLOW data illustrate the problem.
In the trial, led by Eric F Morand and colleagues, the investigational drug enpatoran failed to meet its primary objective. The multiple comparison procedure-modelling analysis of the British Isles Lupus Assessment Group-based Composite Lupus Assessment, or BICLA, dose-response relationship returned a p-value of 0.14, falling short of statistical significance.
However, the correspondence points to a response pattern that does not follow a simple dose-dependent curve. Placebo response was recorded at 39 percent, while the three active dose groups produced responses of 58 percent, 49 percent, and 49 percent, respectively. This non-monotonic pattern, in which the highest response appears at the lowest dose and then plateaus, suggests a pharmacodynamic ceiling rather than an absence of drug activity.
The distinction matters because a composite endpoint like BICLA is designed to capture a broad set of clinical changes across multiple organ systems. If a drug's mechanism acts on a narrower biological pathway, the composite may dilute the signal. The correspondence argues that the WILLOW results should be read as a case of mechanism-endpoint mismatch, not necessarily as evidence that enpatoran is inactive.
Lupus remains one of the most challenging autoimmune diseases to treat, and trial design has long been a source of frustration for researchers and patients alike. The disease's heterogeneous presentation means that composite endpoints must balance sensitivity across different manifestations, sometimes at the cost of detecting targeted biological effects.
The WILLOW trial's Cohort B was designed to test enpatoran, a drug that targets a specific immunological pathway. According to the correspondence, the failure of the primary analysis should not close the book on the compound. Instead, it raises questions about whether the BICLA-based dose-response model was the appropriate tool for measuring the drug's impact.
The authors of the correspondence do not dispute the statistical outcome of the trial. Their argument is interpretive: the non-monotonic response curve is a signal that the drug engaged its target, but that the endpoint may have been insensitive to the resulting biological change. They call for greater attention to matching mechanism with endpoint in future lupus studies.
This is not the first time lupus trials have faced such criticism. The disease's complex immunology and the lack of biomarkers that reliably predict clinical benefit have made it difficult to design trials that satisfy both regulatory requirements and mechanistic curiosity. The WILLOW case adds a concrete example to that ongoing methodological conversation.
For clinicians and researchers, the implication is that negative primary results in lupus should be examined for patterns that might indicate pharmacodynamic effects. A flat dose-response curve suggests no activity; a plateau or non-monotonic curve suggests that the drug is doing something, but that the trial's measurement strategy may not be capturing it.
The correspondence does not propose a specific alternative endpoint for WILLOW, nor does it call for a reanalysis of the trial data. It uses the Cohort B results as a teaching case for the broader field, urging investigators to consider whether their chosen composite endpoints are aligned with the mechanism of the drugs they are testing.
As lupus research continues to attract investment from both industry and academic groups, the debate over trial design is likely to intensify. The WILLOW Cohort B data, with their unusual response pattern, offer a reminder that statistical failure and biological failure are not always the same thing.
15





