# Reposting: Dwi2response Dhollander: extreme SDM tails causing very small CSF candidate pools

**URL:** https://community.mrtrix.org/t/reposting-dwi2response-dhollander-extreme-sdm-tails-causing-very-small-csf-candidate-pools/8836
**Category:** Uncategorized
**Created:** [September 15, 2026, 7:42am UTC](https://community.mrtrix.org/t/reposting-dwi2response-dhollander-extreme-sdm-tails-causing-very-small-csf-candidate-pools/8836 "2026-09-15T07:42:53Z")
**Posts on this page:** 6
**Page:** 1

<div class="post-metadata">

### Author: ![WilliamFCB](https://community.mrtrix.org/letter_avatar_proxy/v4/letter/w/e68b1a/32.png) [@WilliamFCB](https://community.mrtrix.org/u/WilliamFCB)
#### Post date: [September 15, 2026, 7:42am UTC](https://community.mrtrix.org/t/reposting-dwi2response-dhollander-extreme-sdm-tails-causing-very-small-csf-candidate-pools/8836/1 "2026-09-15T07:42:53Z")

</div>

I repost this as the the original post is greyed out and contains " preprocessing error", I did not get any response

Hi MRtrix3 team, @ThijsDhollander @jdtournier

I am investigating reproducible dwi2response dhollander failures in a small number of subjects. After inspecting the retained temporary files, I now seem to have two related failure patterns.

For two subjects, the crude CSF pool is large (~13–14k voxels) (e.g attached CASE 1), but CSF refinement collapses it to only 1–2 voxels. The subsequent CSF response-voxel selection then fails because 10% of the refined pool rounds to zero.

For another subject (attached CASE2), the problem occurs earlier. The safe\_mask contains 402,936 voxels and is partitioned as:

```auto
crude WM: 193,182

crude GM: 209,752

crude CSF: 2

```

The FA image and DWI volumes look visually okay. The two voxels classified as crude CSF have very high SDM values (one around 7), while essentially the entire remaining non-WM pool is classified as crude GM.

From reading the Dhollander implementation, my understanding is that automatic mrthreshold operations are used both when splitting the low-FA/non-WM pool into crude GM versus crude CSF and later during CSF refinement.

It seems that very small number of extreme SDM values can cause the automatic threshold to isolate only the extreme tail, resulting in biologically implausibly small CSF pools even when the underlying DWI and FA images appear otherwise reasonable

I would be very interested in your thoughts on:

. Is this behaviour an expected failure mode of the current Dhollander algorithm?

. Is it correct that a handful of extreme SDM values dominate the automatic mrthreshold split in either the crude GM/CSF separation or the later CSF refinement?

. Is there any built-in safeguard to prevent a tissue candidate pool from collapsing to only one or a few voxels?

. If the extreme-SDM voxels can be shown spatially to be implausible and/or associated with abnormal underlying DWI signal, would excluding those voxels from the response-estimation mask and rerunning otherwise standard Dhollander be a reasonable approach?

. Would it make sense for Dhollander to flag an extreme crude→refined or non-WM→CSF collapse before proceeding to response-voxel selection?

I attached, SDM histograms before and after the relevant splits, crude/refined tissue counts, spatial locations of the extreme-SDM voxels, and their underlying per-volume DWI signal.

Please your response  
BW  
William

[dwi2response.pdf](https://community.mrtrix.org/uploads/short-url/21Md6S5hcQK5PFm9s8MGvke05v7.pdf) (3.4 MB)

---

<div class="post-metadata">

### Author: ![jdtournier](https://community.mrtrix.org/user_avatar/community.mrtrix.org/jdtournier/32/2594_2.png) [@jdtournier](https://community.mrtrix.org/u/jdtournier)
#### Post date: [September 15, 2026, 4:27pm UTC](https://community.mrtrix.org/t/reposting-dwi2response-dhollander-extreme-sdm-tails-causing-very-small-csf-candidate-pools/8836/2 "2026-09-15T16:27:06Z")

</div>

Hi @WilliamFCB,

Thanks for persisting, and sorry for the slow response (I’m finding it increasingly difficult to find time to help out here, unfortunately).

> - Is this behaviour an expected failure mode of the current Dhollander algorithm?

No, not an expected failure mode…

That said, these types of issues may also be related to idiosyncratic features of your data, and it may be worth investigating what is causing these values in the first place. Maybe you could provide details of your acquisition, and maybe screenshots of the MD (ADC) maps computed from these data (including the problematic voxels)? I note there are a few screenshots in your attached PDF (very detailed, thank you!), but I don’t think any of them show the MD?

> - Is it correct that a handful of extreme SDM values dominate the automatic mrthreshold split in either the crude GM/CSF separation or the later CSF refinement?

Not correct, but not unexpected given how the algorithm works. When the threshold is not specified, `mrthreshold` uses the method described in this paper:

> Ridgway, G. R.; Omar, R.; Ourselin, S.; Hill, D. L.; Warren, J. D. & Fox, N. C. Issues with threshold masking in voxel-based morphometry of atrophied brains. NeuroImage, 2009, 44, 99-111

And I can see how a few abnormally large values can affect the results.

> - Is there any built-in safeguard to prevent a tissue candidate pool from collapsing to only one or a few voxels?

No, none that I know of…

> - If the extreme-SDM voxels can be shown spatially to be implausible and/or associated with abnormal underlying DWI signal, would excluding those voxels from the response-estimation mask and rerunning otherwise standard Dhollander be a reasonable approach?

Yes, I think that would be a simple way to work around that problem. Anything else is likely to require modifications to the algorithm itself, which is going to take some time and effort (which, as mentioned earlier, is in short supply).

> - Would it make sense for Dhollander to flag an extreme crude→refined or non-WM→CSF collapse before proceeding to response-voxel selection?

Yes, that would indeed be a simple diagnostic to add, though not an actual fix. Maybe we could check whether the count in the mask is at least some number (10, say), and if not, re-run `mrthreshold` with the `-top 10` option to request the top 10 voxels instead? That should be a relatively simple change to make to the script… Maybe you’d like to give that a shot?

All the best,  
Donald

---

<div class="post-metadata">

### Author: ![WilliamFCB](https://community.mrtrix.org/letter_avatar_proxy/v4/letter/w/e68b1a/32.png) [@WilliamFCB](https://community.mrtrix.org/u/WilliamFCB)
#### Post date: [September 17, 2026, 11:22am UTC](https://community.mrtrix.org/t/reposting-dwi2response-dhollander-extreme-sdm-tails-causing-very-small-csf-candidate-pools/8836/3 "2026-09-17T11:22:31Z")

</div>

Hi Donald,

Thanks for your reply. We looked more closely at the failed subjects, and I think we’ve narrowed the problem down quite a bit.

In all three failed subjects, safe\_fa.mif contained only a very small number of non-finite FA voxels: 1, 3 and 2 voxels, respectively. Looking back at the original DWI signal at these six locations, all showed a very similar pattern: 76–79 of the 80 b≈1000 measurements were ≤0, while the low-b signal was much better preserved. This produced a very large apparent signal decay, with a small remaining positive sign, and large SDM.

Our understanding of the 3.0.8 dhollander.py code is that safe\_mask is created before tensor fitting by excluding voxels with non-finite or non-positive shell-mean signal or SDM values. However, this does not exclude voxels with an extremely high but still finite SD, as in our 3 cases.

This seems to explain why the affected voxels survived this step in our subjects. Although 76–79 of their 80 b≈1000 measurements were ≤0, negative measurements are clipped to zero before calculating the shell mean. A small remaining positive signal was therefore sufficient to produce a finite positive shell mean and an extreme but finite SDM. The subsequent tensor fit then produced non-finite FA at these locations. We could not find that a subsequent finite-FA check was performed before crude tissue classification. Consequently, non-finite voxels could enter the non-WM pool.

Of five “control” subjects for whom Dhollander completed normally, none had non-finite FA within safe\_mask, and their refined-CSF pools contained 4,279–10,679 voxels.

We reran the three failed subjects after removing the non-finite safe\_fa voxels from the response mask. All three completed normally: refined CSF increased to 4,828, 4,562 and 7,090 voxels, with final CSF response selections of 483, 456 and 709 voxels. The refined-CSF SDM distributions are in the same range as the successful controls.

We wondered whether a small check after safe\_fa is calculated might be more direct: i.e., keep the existing thresholds and refinement logic, but exclude voxels with non-finite safe\_fa values from the mask used for the subsequent crude tissue classification. Setting non-finite FA to zero would still place those “abnormal” voxels on the non-WM side of the FA threshold. Thus, better to exclude them from that classification step altogether.

Would filtering out non-finite safe\_fa voxels at that point be a sensible place to handle this, or is there a reason non-finite voxels are allowed to remain in safe\_mask?

We haven’t tested the “-top 10” fallback you suggested. Probably, it wouldn’t address the underlying issue in these cases and would keep “abnormal” voxels with extreme SDM in the pool, even though it may prevent a final 10% selection from rounding to zero.

**One additional observation:** in the successful reruns and controls, we noticed that crude CSF is entirely FA≤0.2 as expected, but finite FA\>0.2 voxels are reintroduced when high-SDM crude-WM voxels are added to the CSF refinement candidates. A small number (~2.3 %) then remain in refined CSF and the final response selection. Is this intended behaviour? There seems to be enough “real” CSF voxels to work with.

I attached a short summary with the failed runs, the three reruns, five successful controls, SDM/FA distributions and the underlying signal at the non-finite-FA voxels.

BW  
[MRtrix\_Dhollander\_nonfinite\_FA.pdf](https://community.mrtrix.org/uploads/short-url/xzXft4Zrs5Jim73hNpVAGlQy85m.pdf) (594.0 KB)

William

---

<div class="post-metadata">

### Author: ![jdtournier](https://community.mrtrix.org/user_avatar/community.mrtrix.org/jdtournier/32/2594_2.png) [@jdtournier](https://community.mrtrix.org/u/jdtournier)
#### Post date: [September 17, 2026, 3:38pm UTC](https://community.mrtrix.org/t/reposting-dwi2response-dhollander-extreme-sdm-tails-causing-very-small-csf-candidate-pools/8836/4 "2026-09-17T15:38:45Z")

</div>

Thanks for the comprehensive report! It’ll take me a while to digest, but hopefully your suggestion of culling out non-finite FA voxels is something we can implement relatively quickly. I’ll see if I have time to look into it in the near future.

BTW, don’t hesitate to have a go yourself and make a pull request on [GitHub](https://github.com/MRtrix3/mrtrix3/)! Contributions are always welcome (even if we struggle to find time to review the larger ones…).

---

<div class="post-metadata">

### Author: ![WilliamFCB](https://community.mrtrix.org/letter_avatar_proxy/v4/letter/w/e68b1a/32.png) [@WilliamFCB](https://community.mrtrix.org/u/WilliamFCB)
#### Post date: [September 19, 2026, 6:09am UTC](https://community.mrtrix.org/t/reposting-dwi2response-dhollander-extreme-sdm-tails-causing-very-small-csf-candidate-pools/8836/5 "2026-09-19T06:09:12Z")

</div>

Hi Donald,

I will make minimal changes tot the Dhollander code and test a “finite\_fa\_safe\_mask” on my test data. I am abroad next week, so I will do it after I get back

Cheers William

---

<div class="post-metadata">

### Author: ![WilliamFCB](https://community.mrtrix.org/letter_avatar_proxy/v4/letter/w/e68b1a/32.png) [@WilliamFCB](https://community.mrtrix.org/u/WilliamFCB)
#### Post date: [September 28, 2026, 6:03am UTC](https://community.mrtrix.org/t/reposting-dwi2response-dhollander-extreme-sdm-tails-causing-very-small-csf-candidate-pools/8836/6 "2026-09-28T06:03:35Z")

</div>

Hi Donald,

We minimally adjusted the code. We created a finite\_fa\_safe\_mask.mif containing only voxels from safe\_mask.mif with finite FA. This mask is then used for the existing crude WM / crude non-WM classification. Everything downstream is unchanged.

I tested the patched Dhollander on the three subjects that previously failed:  
CASE01: 1 non-finite FA voxel excluded; Dhollander completed successfully  
CASE02: 3 non-finite FA voxels excluded; Dhollander completed successfully  
CASE03: 2 non-finite FA voxels excluded; Dhollander completed successfully

So in all three cases, excluding exactly the non-finite-FA voxels at this point was sufficient for the normal Dhollander workflow to complete.

I also successfully ran control subjects. They contained no non-finite FA voxels, so finite\_fa\_safe\_mask was identical to safe\_mask. The resulting responses were essentially unchanged; the only differences I saw were numerical rounding at approximately the 14th–15th decimal place.

I am not that git savvy, so below the lines of code we changed:

```plaintext
- run.command('mrcalc safe_mask.mif safe_fa.mif 0 -if ' + str(app.ARGS.fa) + ' -gt crude_wm.mif -datatype bit', show=False)
- run.command('mrcalc crude_wm.mif 0 safe_mask.mif -if _crudenonwm.mif -datatype bit', show=False)
+ run.command('mrcalc safe_mask.mif safe_fa.mif -finite -mult finite_fa_safe_mask.mif -datatype bit', show=False)
+ run.command('mrcalc finite_fa_safe_mask.mif safe_fa.mif 0 -if ' + str(app.ARGS.fa) + ' -gt crude_wm.mif -datatype bit', show=False)
+ run.command('mrcalc crude_wm.mif 0 finite_fa_safe_mask.mif -if _crudenonwm.mif -datatype bit', show=False)

```

I will work with the local patched version for now. Please let me know if/when the changes make it to the mrtrix3 distribution

Any thought on my previous observation?  
" **One additional observation:** in the successful reruns and controls, we noticed that crude CSF is entirely FA≤0.2 as expected, but finite FA\>0.2 voxels are reintroduced when high-SDM crude-WM voxels are added to the CSF refinement candidates. A small number (~2.3 %) then remain in refined CSF and the final response selection. Is this intended behaviour? There seems to be enough “real” CSF voxels to work with."

Best,  
William
