Abstract
Neuroscience has been diagnosed with a pervasive lack of statistical power and, in turn, reliability. One remedy proposed is a massive increase of typical sample sizes. Parts of the neuroimaging community have embraced this recommendation and actively push for a reallocation of resources toward fewer but larger studies. This is especially true for neuroimaging studies focusing on individual differences to test brain–behavior correlations. Here, I argue for a more efficient solution. Ad hoc simulations show that statistical power crucially depends on the choice of behavioral and neural measures, as well as on sampling strategy. Specifically, behavioral prescreening and the selection of extreme groups can ascertain a high degree of robust in-sample variance. Due to the low cost of behavioral testing compared to neuroimaging, this is a more efficient way of increasing power. For example, prescreening can achieve the power boost afforded by an increase of sample sizes from n = 30 to n = 100 at ∼5% of the cost. This perspective article briefly presents simulations yielding these results, discusses the strengths and limitations of prescreening and addresses some potential counter-arguments. Researchers can use the accompanying online code to simulate the expected power boost of prescreening for their own studies.
Introduction
Recent estimates show that the statistical power of typical studies in neuroscience is inadequately low (). Low power and a publication bias for significant results lead to low replicability of published findings. Many researchers, journals and funding agencies are acutely aware of the problem and discuss a range of potential remedies, including the publication of data and analysis code (), preregistration () and an increase of typical sample sizes ().
In parallel to this, cognitive neuroscience is gaining interest in individual differences revealing brain–behavior correlations (; ). This trend includes subfields like visual neuroscience (; ; ), which traditionally have treated such differences as ‘noise’ (; ; ). Studies investigating individual differences typically require larger sample sizes, and in light of the replicability debate, recently proposed a new standard of ‘n > 100.’
Here, I argue for a more efficient way of increasing the power to detect brain–behavior correlations. Simulations show that adequate power can be achieved with a prescreening approach, in which researchers test a larger sample behaviorally and selectively sample extreme-groups for brain scanning. Compared to the proposed increase in sample size, this approach typically enhances power at a fraction of the cost. Additionally, prescreening can ensure the reliability of measures, which is limiting observable effect sizes (see below).
Many others have discussed the importance of reliable measures (see below), as well as the advantages and limitations of prescreening (; ; ; ). This perspective does not aim to make original points about either issue as such. Instead, it aims to highlight their special importance in studies using expensive techniques like MEG or MRI to study brain–behavior correlations. This boundary condition renders prescreening an efficient way to well-powered studies.
Ensuring Reliability
The reliability of inter-individual differences in a given behavioral or neural measure depends on two sources. It increases with true between-subject variance of the measured trait and decreases with measurement error. Reliability can be estimated as the consistency between trials, items or parallel forms, or as test–retest reliability. As has been pointed out by others (), a given measure can robustly detect an effect at the group level, without reliably capturing individual differences. Personality and intelligence research have a history of developing reliable measures of behavioral variance (with the search for robust neural correlates of these measures proving more difficult; ; ). But as the investigation of brain–behavior correlations is adopted in other fields, the reliability of behavioral measures is often unknown a priori. Prescreening can be used to ensure and quantify this quality.
For example, and recently argued that face inversion effects are driven by retinotopic tuning biases. A potential way of testing this hypothesis would be to probe a correlation between the corresponding neural and behavioral effects across observers. However, these authors did not use this type of strategy, because their measures did not reliably capture individual differences. Figure 1A shows the individual magnitude of the Thatcher illusion () for 36 observers in . The illusion was highly robust – every observer showed the effect for both, odd and even trials (all data are in the upper right quadrant). At the same time, the inter-individual variance was highly inconsistent (r = 0.20). That is, even though any two observers consistently show an effect greater than zero, the difference between them is not reliable.1
FIGURE 1
Compare this to Figure 1B, showing individual differences in proneness to the sound-induced flash illusion (
Typically, a priori knowledge about the reliability of brain measures is rather limited. The reliability of neural measures depends on a range of factors including participant state, imaging method, the amount of data collected per participant, hardware, acquisition parameters, experimental design, preprocessing, and analysis pipelines. A thorough review of these factors is beyond the scope of the current manuscript [for an overview regarding fMRI, including pointers to optimization tools see (
The importance of reliable measures has been pointed out before (
Figure 1C illustrates the effect of this for different effect sizes and levels of reliability (). Researchers undertaking power calculations have to decide for a minimum effect size their study should be sensitive for. They may decide that a biologically meaningful effect implies a minimum of ∼9% shared variance and therefore aim for adequate power (>85%) to detect brain–behavior correlations > = 0.3. However, if the measures used have limited reliability, this has to be taken into account. Adequate power for an observable effect size ro = 0.3 will correspond only to (much) stronger biological effects rh if they are attenuated by unreliable measures (the horizontal dashed line in Figure 1C). Even a relatively moderate lack of reliability will result in a noticeable drop in power. Using measures with a reliability of 0.7, sensitivity for observable effects ro = 0.3 would translate to biological effects rh > 0.43. Conversely, preserving adequate power for ‘true’ effects rh > 0.3 would require sensitivity for observable effect sizes ro > 0.21 (vertical dashed line in Figure 1C). This implies an approximate doubling of the required sample size from n = 97 to n = 201.
Wherever possible, an investment in reliable measures seems more efficient than bringing a large sample with unreliable measures to the scanner. For novel measures, reliability should be quantified and reported (
Sampling Selectively
The power to detect covariance depends on (true) in-sample variance. Prescreening allows maximizing variance through selective sampling. Here, I will briefly present the results of ad hoc simulations of this effect. The code is available at https://osf.io/hjdcf/ and interested readers can turn there to find more details and adjust parameters for their own power calculations.
To simulate the effect of a prescreening strategy, 10 bivariate populations were created by drawing 107 normally distributed random ‘behavioral’ values x ∼ N(0,1). Corresponding ‘brain’ values (y) were simulated based on the normally distributed random variable e ∼N(0,1) and a defined observable brain–behavior correlation ro (specific for each population and ranging from 0 to 0.9), such that
Results of y were only accepted if they correlated with x by ro within a tolerated error margin of 0.01 (otherwise the procedure was repeated). From each of the resulting populations, 10.000 random samples with n observations were drawn with replacement (n ranging from 20 to 120 in steps of 10). For each level of ro and n, Power was estimated as the fraction of the 10.000 samples showing a significant brain–behavior (Pearson) correlation (P < 0.05). Figure 2A shows the relationship between power, sample size and ro. In line with analytic predictions and recent recommendations (
FIGURE 2

Selective sampling. Left hand plots in black ink show power simulations without prescreening; right hand plots in blue show corresponding results for prescreening with selective sampling of extreme groups. Selective samples were drawn based on behavioral measures from a prescreening group six times the size of n (extreme groups for B,D, even sampling for G). In each panel, ink saturation indicates the simulated effect size on the population level, as shown in the inset (observable correlation ro). Red ink indicates simulation results for zero-effects (ro = 0). All x axes indicate sample size n (10.000 random samples drawn for each level). (A,B) Shows power, i.e., the fraction of samples showing a statistically significant correlation (P < 0.05). (C,D) Shows corresponding effect sizes (mean of observed significant correlations). (E) Shows the (inverse) precision of effect size estimates (width of 95% confidence interval) as a function of observed effect size and sample size. (F,G) Shows the accuracy of a model selection procedure aiming to distinguish between linear and non-linear relationships based on simulated data (see main text for details). (H) Shows the power boost afforded by different prescreening factors, for two different observable effect sizes (ro = 0.3/0.4, as shown). The prescreening factor corresponds to the size of the prescreening sample divided by the final n, as shown in the legend to the right.
Figure 2B shows the results of the same simulation incorporating prescreening. Here, for each sample size, behavioral measures (x) were first drawn for a larger prescreening sample, six times the size of n. Brain–behavior (Pearson) correlations were only probed in a subsample, for which n participants with extreme behavioral values were selected from the prescreening sample (n/2 participants with the lowest and highest prescreened values in x, respectively). This strategy resulted in a remarkable power boost (Figure 2B). With prescreening, good power (>85%) to detect ro > 0.3 could be achieved with an in-scanner sample size of n = 30. Importantly, this power boost was limited to true effects. For ro = 0 the nominal false alarm rate of P = 0.05 was preserved. Figure 2H shows corresponding results for further prescreening factors (i.e., multiples of n other than six).
The power boost afforded by prescreening does not come for free. To achieve comparable power to n = 100 at n = 30, the simulated researcher prescreened 180 participants. But the higher cost of neural compared to behavioral testing should typically render this a sensible choice. These factors can vary, but it is worth considering an example.
In the author’s experience, many behavioral experiments can be done in an hour, and the booking time for a typical fMRI session is 2 h. The standard fee for participant reimbursement at his current institution is 8€/h and that for booking a research MRI machine in Germany is 150€/h2. Assuming these numbers, we can calculate the price tag of improving a traditional, low-powered study (n = 30) to a well-powered one using different strategies. Prescreening in this example would require additional behavioral testing of 150 participants, which works out to additional costs of 150∗1h∗8€/h = 1.200€. A similar level of power could be achieved by testing and scanning 70 additional participants, which would work out to 70∗(1h∗8€/h + 2h∗(8€/h + 150€/h)) = 22.680€. Remarkably, prescreening achieves a similar power boost as the recommended increase in sample size at ∼5% of the cost.
I expect this estimate to be a conservative one. The assumed cost of scanning is a fraction of that charged in many centers outside Germany. The example also left aside all staff costs. A single student can typically do behavioral testing, while most neuroimaging facilities will require at least two professionally trained operators. Likewise, the analysis of neuroimaging data requires specialized training and can be more time consuming than that of behavioral experiments. Researchers and funding agencies are encouraged to do their own calculations, adjusting parameters as needed.
Note that error in this simulation was captured by the error term , which negatively scales with ro (Eq. 2). The observable correlation ro in turn depends on the shared variance between ‘true’ behavioral and neural traits, rh, as well as on measurement errors attenuating this relationship (Eq. 1). Expanding ro to directly enter measurement error into the simulation yields identical results (also see section “Counterargument 3: Real Data May Be ‘Nastier’ Than Simulations” below). But specifying ro in a separate, first step seems closer to practical purposes.
To consider an example: A researcher hypothesizes that retinotopically determined V1 surface area negatively correlates with individual proneness to the sound induced flash illusion (c.f.
Caveats and Counterarguments
Conterargument 1: Sampling Extreme Groups Will Yield Inflated Estimates of Effect Sizes
Yes. Researchers applying prescreening should be aware of this and highlight it in their publications. The application of correlation measures across extreme groups is well-established (
Figures 2C,D shows average effect size estimates with and without prescreening for the power simulation shown in Figures 2A,B. Effect size estimates are inflated in both cases, because they are based on significant results only [the simulation assumes a ‘ file drawer problem’ (
However, inflation is limited to true effects; for ro = 0, the average effect size stays zero. Prescreening preserves a low false positive rate, while affording a high probability of detecting true effects. Just as for increased sample sizes, this results in a high positive predictive value of significant findings. That is, a significant result has a higher probability of reflecting a true effect when the study was well-powered – regardless of whether that power was achieved through prescreening or increased sample sizes.
Selective sampling is an efficient strategy for detecting (or rejecting) brain–behavior correlations. It is not suitable for precise estimates of their size or the parameters of a predictive model (c.f.
Finally, the description of a population parameter requires a well-defined population (
Counterargument 2: Sampling Extreme Groups Can Conceal Non-linear Brain–Behavior Relations
Yes. The extreme-group strategy proposed here aims at detecting (quasi-)linear relationships. Researchers aiming to compare different models should optimize their sampling strategy accordingly.
Figures 2F,G show the results of a simulation of ‘true’ linear and non-linear brain–behavior relationships. This simulation followed three steps. First, it drew random behavioral data (x) from a normal distribution x ∼N(0,1). Second, idealized brain predictions (yh) were generated, 50% of which perfectly corresponded to the model:
and 50% of perfectly corresponded to the model
Importantly, a third step added brain measurement noise, which was manipulated to be comparable for both models and to the power calculations above. Specifically, brain measures (y) were simulated as random data linearly correlated to the respective model predictions (yh) with ro. In this way, 10.000 samples were drawn for both models at each level of ro and sample sizes ranging from n = 20 to n = 120.
For each sample, the simulation aimed to distinguish linear from non-linear relationships by fitting 1st and 2nd order polynomials. The ‘wining’ model for each sample was chosen based on the Akaike Information Criterion (
To evaluate the usefulness of prescreening in such a scenario, the simulation was repeated for behavioral prescreening with n∗6. The sampling strategy was adjusted to the needs of model comparison. Instead of sampling the tails of the prescreening sample, the algorithm aimed at choosing a subsample that covered the behavioral range of the prescreening sample as evenly as possible (see online code for details). Figure 2G shows that this strategy significantly enhanced the sensitivity to discriminate non-linear from linear relationships.
Counterargument 3: Real Data May Be ‘Nastier’ Than Simulations
Yes. The simulations presented here assume normally distributed data and random errors. Real data may be less well-behaved. However, it is not clear a priori that this would pose more of a problem for the prescreening compared to a full-sample approach. Moreover, prescreening samples enable informed hypotheses regarding potentially problematic aspects of the data.
A particularly relevant example of problematic data is that of heteroscedastic measurement error, scaling with the latent variable. In this scenario, extreme groups will be affected by particularly low and high measurement errors, respectively. Importantly, supplementary simulations confirm the robustness of prescreening in this situation. Prescreening and selective sampling preserved nominal false positive rates and a strong power boost, even for a population with strongly heteroscedastic measurement error (Supplementary Data and Supplementary Figure 1).
Counterargument 4: Extreme Groups May Be Special
Yes, but it is important to spell out what that means.
It may refer to the relationship of a behavioral and a neural trait not following a uniform, linear model across the entire distribution. In such cases, the underlying linear model is inappropriate, regardless of the sampling approach. See Counterargument 2 for prescreening in the context of non-linear model comparisons.
Alternatively, the argument may refer to measurement error. Even random measurement errors will correlate with observed values. That is, extreme measurements will partly reflect extreme errors (and more so for less reliable measures). This biased error sampling causes regression to the mean (
Finally, confounding factors may be especially pronounced in behavioral extreme groups. Statistical control for confounding factors can be difficult if they have to be estimated with noisy measures (
Counterargument 5: Prescreening Should Be Combined With Large Samples Rather Than Pitted Against Them
Not necessarily.
Researchers interested in precise effect size estimates will indeed need four-figure sample sizes, unless the effects they study are unusually large. However, the aim of many brain–behavior studies is arguably more humble. Researchers often are interested in testing the hypothesis that there is some relationship between a given neural and a behavioral measure which is captured well-enough by a linear model to be relevant in the context of their theory (e.g., ro > 0.3). Testing this hypothesis can be a valuable first step, even though (for affirmative cases) it certainly should not be the last (
Prescreening can ensure adequate power with small sample sizes in the scanner, even for moderate effects (e.g., >85% power for ro > 0.3 at n = 30). Researchers combining this approach with individually consistent measures can be confident in the detection power and replicability7 of their studies (Figure 2B). The savings afforded by this strategy relative to a blanket increase of sample sizes should routinely be around a factor 20 or higher. Given the limited resources available to neuroscience, these savings would translate to a larger number of well-powered detection studies.
Counterargument 6: Prescreening Precludes the Reuse of Data for Unrelated Research Questions
Yes (usually). Neuroimaging experiments can produce data that are highly specific to an underlying research question, like the retinotopic specificity of visual cortex responses to illusory contours (
Not in the author’s opinion. The best sources of ‘general purpose’ data are public datasets like the aforementioned Human Connectome Project. These are truly large scale, representative and include extensive test batteries. This will typically not be the case for data from individual experiments. Consider a researcher investigating the relationship between V1 surface area and the individual strength of contextual size illusions (
Counterargument 7: Prescreening Sometimes Is Not Feasible
Yes. For instance, it can be more efficient to scan every participant if the recruitment process itself is costly (as for some special populations). In general, the cost ratio of behavioral and neural testing will vary with the type of behavioral measure. The example calculation given above assumed a simple, lab-based psychophysics or eyetracking experiment. The potential advantage of prescreening will be even more pronounced for questionnaires, or tasks that can be completed online. At the other end of the spectrum are resource-intensive behavioral tasks, like reverse correlation techniques requiring 10s of 1000s of trials (
Furthermore, for some behavioral tasks it may be impossible to estimate the reliability of between subject variance, for instance because they require participants to be naïve and cannot be repeated (like some measures of perceptual learning). However, such scenarios will typically call for an entirely different research design. Measures that cannot be tested for the reliability of between-subject variance are unsuitable for brain–behavior correlations as such, not just for prescreening. Group level effects can be robust independently of this (Figure 1A) and should be the variable of interest in this type of situation.
Conclusion
Like all of science, studies aiming at the detection of brain–behavior correlations depend on well-powered experiments, yielding replicable results. This perspective highlighted how this can be achieved with relatively small sample sizes in the scanner. Behavioral prescreening can achieve a power boost comparable to larger sample sizes at a fraction of the cost and without inflating the false positive rate.
Researchers investigating brain–behavior correlations should base their power calculations on a broader basis than sample size alone. The simulation code accompanying this perspective can be used to incorporate the boost afforded by prescreening. Researchers should also pay attention to the reliability of measures and adjust their power estimates for attenuation. Similarly, reviewers, editors and funding agencies should refrain from a simple but false heuristic of large in-scanner samples as a necessary or sufficient criterion for adequate detection power.
Statements
Author contributions
BdH: coded simulations, prepared figures, and wrote the paper.
Funding
This work was supported by a JUST’US fellowship from Justus-Liebig-Universität Gießen.
Acknowledgments
I thank Denis-Alexander Engemann, Bertram Walter, Axel Schäfer, Johannes Haushofer, Will Harrison, Bianca Wittmann, Huseyin Boyaci, and Anke-Marit Albers for discussions of related subjects and/or earlier versions of this manuscript. I am solely responsible for any remaining errors and all opinions expressed.
Conflict of interest
The author declares that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
Supplementary material
The Supplementary Material for this article can be found online at: https://www.frontiersin.org/articles/10.3389/fnhum.2018.00421/full#supplementary-material
Footnotes
1.^Note that this may partly be due to ceiling effects and a limited number of trials. Individual face inversion effects in other designs may vary more consistently (
2.^http://www.dfg.de/formulare/55_04/55_04_de.pdf
5.^http://www.humanconnectomeproject.org/
6.^http://www.ukbiobank.ac.uk/
7.^Note that replicability here refers to the detection question: experiments will give reliable answers to the question whether there is an effect whenever there is no effect (repetitions of the experiment will yield mostly negative results) or a true effect of a magnitude they are well-powered for (repetitions of the experiment will mostly yield significant results). Replicability in a stricter sense of near-identical effect size estimates is much harder to achieve (3.1). Both concepts are linked in the sense that more precise effect size estimates imply power for smaller effect sizes. That is, more precise effect size estimates reduce the range of true effects the experiment is not well-powered for.
References
1
AbrahamsN. M.AlfE. F. (1978). Relative costs and statistical power in the extreme groups approach.Psychometrika4311–17. 10.1007/BF02294085
2
AkaikeH. (1974). A new look at the statistical model identification.IEEE Trans. Autom. Control19716–723. 10.1109/TAC.1974.1100705
3
AlfE. F.AbrahamsN. M. (1975). The use of extreme groups in assessing relationships.Psychometrika40563–572. 10.1007/BF02291557
4
BarnettA. G.van der PolsJ. C.DobsonA. J. (2004). Regression to the mean: what it is and how to deal with it.Int. J. Epidemiol.34215–220. 10.1093/ije/dyh299
5
BennettC. M.MillerM. B. (2010). How reliable are the results from functional magnetic resonance imaging?Ann. N. Y. Acad. Sci.1191133–155. 10.1111/j.1749-6632.2010.05446.x
6
BensonN.JamisonK. W.ArcaroM. J.VuA.GlasserM. F.CoalsonT. S.et al (2018). The HCP 7T retinotopy dataset: description and pRF analysis.bioRxiv [Preprint]. 10.1101/308247
7
BrandtD. J.SommerJ.KrachS.BedenbenderJ.KircherT.PaulusF. M.et al (2013). Test-retest reliability of fMRI brain activity during memory encoding.Front. Psychiatry4:163. 10.3389/fpsyt.2013.00163
8
ButtonK. S.IoannidisJ. P. A.MokryszC.NosekB. A.FlintJ.RobinsonE. S. J.et al (2013). Power failure: why small sample size undermines the reliability of neuroscience.Nat. Rev. Neurosci.14365–376. 10.1038/nrn3475
9
BzdokD.EngemannD.-A.GriselO.VaroquauxG.ThirionB. (2018). Prediction and inference diverge in biomedicine: simulations and real-world data.bioRxiv [Preprint]. 10.1101/327437
10
ChambersC. D. (2013). Registered reports: a new publishing initiative at cortex.Cortex49609–610.
11
CharestI.KriegeskorteN. (2015). The brain of the beholder: honouring individual representational idiosyncrasies.Lang. Cogn. Neurosci.30367–379. 10.1080/23273798.2014.1002505
12
de HaasB.KanaiR.JalkanenL.ReesG. (2012). Grey matter volume in early human visual cortex predicts proneness to the sound-induced flash illusion.Proc. Biol. Sci.2794955–4961. 10.1098/rspb.2012.2132
13
de HaasB.SchwarzkopfD. S. (2018a). Feature-location effects in the Thatcher illusion.J. Vis.18:16. 10.1167/18.4.16
14
de HaasB.SchwarzkopfD. S. (2018b). Spatially selective responses to Kanizsa and occlusion stimuli in human visual cortex.Sci. Rep.8:611. 10.1038/s41598-017-19121-z
15
de HaasB.SchwarzkopfD. S.AlvarezI.LawsonR. P.HenrikssonL.KriegeskorteN.et al (2016). Perception and processing of faces in the human brain is tuned to typical feature locations.J. Neurosci.369289–9302. 10.1523/JNEUROSCI.4131-14.2016
16
DuboisJ.AdolphsR. (2016). Building a science of individual differences from fMRI.Trends Cogn. Sci.20425–443. 10.1016/j.tics.2016.03.014
17
DuboisJ.GaldiP.HanY.PaulL. K.AdolphsR. (2018). Resting-state functional brain connectivity best predicts the personality dimension of openness to experience.bioRxiv [Preprint]. 10.1101/215129
18
GençE.BergmannJ.SingerW.KohlerA. (2015). Surface area of early visual cortex predicts individual speed of traveling waves during binocular rivalry.Cereb. Cortex251499–1508. 10.1093/cercor/bht342
19
GosselinF.SchynsP. G. (2003). Superstitious perceptions reveal properties of internal representations.Psychol. Sci.14505–509. 10.1111/1467-9280.03452
20
HedgeC.PowellG.SumnerP. (2018). The reliability paradox: why robust cognitive tasks do not produce reliable individual differences.Behav. Res. Methods501166–1186. 10.3758/s13428-017-0935-1
21
HenrichJ.HeineS. J.NorenzayanA. (2010). The weirdest people in the world?Behav. Brain Sci.3361–83. 10.1017/S0140525X0999152X
22
KanaiR.ReesG. (2011). The structural basis of inter-individual differences in human behaviour and cognition.Nat. Rev..Neurosci.12231–242. 10.1038/nrn3000
23
MadanC. R.KensingerE. A. (2017). Test–retest reliability of brain morphology estimates.Brain Inform.4107–121. 10.1007/s40708-016-0060-4
24
Martín-BuroM. C.GarcésP.MaestúF. (2016). Test-retest reliability of resting-state magnetoencephalography power in sensor and source space.Hum. Brain Mapp.37179–190. 10.1002/hbm.23027
25
McEvoyL. K.SmithM. E.GevinsA. (2000). Test-retest reliability of cognitive EEG.Clin. Neurophysiol.111457–463.
26
MollonJ. D.BostenJ. M.PeterzellD. H.WebsterM. A. (2017). Individual differences in visual science: what can be learned and what is good experimental practice?Vis. Res.1414–15. 10.1016/j.visres.2017.11.001
27
MoutsianaC.De HaasB.PapageorgiouA.Van DijkJ. A.BalrajA.GreenwoodJ. A.et al (2016). Cortical idiosyncrasies predict the perception of object size.Nat. Commun.7:12110. 10.1038/ncomms12110
28
PernetC.PolineJ.-B. (2015). Improving functional magnetic resonance imaging reproducibility.Gigascience4:15. 10.1186/s13742-015-0055-8
29
PeterzellD. H.KennedyJ. F. (2016). “Discovering sensory processes using individual differences: a review and factor analytic manifesto”, inProceeding of the Electronic Imaging: Human Vision and Electronic Imaging, Vol. 11, 1–11. 10.2352/ISSN.2470-1173.2016.16.HVEI-112
30
PlichtaM. M.SchwarzA. J.GrimmO.MorgenK.MierD.HaddadL.et al (2012). Test–retest reliability of evoked BOLD signals from a cognitive–emotive fMRI test battery.Neuroimage601746–1758. 10.1016/j.neuroimage.2012.01.129
31
PreacherK. J. (2015). “Extreme groups designs,” inThe Encyclopedia of Clinical PsychologyVol. 2edsCautinR. L.LilienfeldS. O.CautinR. L.LilienfeldS. O. (Hoboken, NJ: John Wiley & Sons, Inc.), 1189–1192.
32
PreacherK. J.RuckerD. D.MacCallumR. C.NicewanderW. A. (2005). Use of the extreme groups approach: a critical reexamination and new recommendations.Psychol. Methods10178–192. 10.1037/1082-989X.10.2.178
33
RezlescuC.SusiloT.WilmerJ. B.CaramazzaA. (2017). The inversion, part-whole, and composite effects reflect distinct perceptual mechanisms with varied relationships to face recognition.J. Exp. Psychol. Hum. Percept. Perform.431961–1973. 10.1037/xhp0000400
34
SchwarzkopfD. S.ReesG. (2013). Subjective size perception depends on central visual cortical magnification in human V1.PLoS One8:e60550. 10.1371/journal.pone.0060550
35
SchwarzkopfD. S.SongC.ReesG. (2011). The surface area of human V1 predicts the subjective experience of object size.Nat. Neurosci.1428–30. 10.1038/nn.2706
36
ShamsL.KamitaniY.ShimojoS. (2000). Illusions: what you see is what you hear.Nature408:788.
37
SimonsohnU.NelsonL. D.SimmonsJ. P. (2014). P-curve: a key to the file-drawer.J. Exp. Psychol. Gen.143534–547. 10.1037/a0033242
38
SmithP. L.LittleD. R. (2018). Small is beautiful: in defense of the small-N design.Psychon. Bull. Rev.1–19. 10.3758/s13423-018-1451-8
39
SpearmanC. (1904). The proof and measurement of association between two things.Am. J. Psychol.15:72. 10.2307/1412159
40
TermenonM.JaillardA.Delon-MartinC.AchardS. (2016). Reliability of graph analysis of resting state fMRI using test-retest dataset from the human connectome project.Neuroimage142172–187. 10.1016/j.neuroimage.2016.05.062
41
ThompsonP. (1980). Margaret Thatcher: a new illusion.Perception9483–484.
42
van DijkJ. A.de HaasB.MoutsianaC.SchwarzkopfD. S. (2016). Intersession reliability of population receptive field estimates.Neuroimage143293–303. 10.1016/j.neuroimage.2016.09.013
43
VulE.HarrisC.WinkielmanP.PashlerH. (2009). Puzzlingly high correlations in fMRI studies of emotion.Pers. Soc. Cogn. Perspect. Psychol. Sci.4274–290. 10.1111/j.1745-6924.2009.01125.x
44
WestfallJ.YarkoniT. (2016). Statistically controlling for confounding constructs is harder than you think.PLoS One11:e0152719. 10.1371/journal.pone.0152719
45
WilmerJ. B. (2008). How to use individual differences to isolate functional organization, biology, and utility of visual functions; with illustrative proposals for stereopsis.Spat. Vis.21561–579. 10.1163/156856808786451408
46
YarkoniT. (2015). “Neurobiological substrates of personality: a critical overview,” inAPA Handbook of Personality and Social Psychology, Personality Processes and Individual DifferencesVol. 4edsMikulincerM.ShaverP. R.SimpsonJ. A.DovidioJ. F. (Washington, DC: American Psychological Association), 61–83. 10.1037/14343-003
47
YovelG.WilmerJ.DuchaineB. (2014). What can individual differences reveal about face processing?Front. Hum. Neurosci.8:562. 10.3389/fnhum.2014.00562
Summary
Keywords
power, replication, individual differences, fMRI, MEG
Citation
de Haas B (2018) How to Enhance the Power to Detect Brain–Behavior Correlations With Limited Resources. Front. Hum. Neurosci. 12:421. doi: 10.3389/fnhum.2018.00421
Received
14 June 2018
Accepted
28 September 2018
Published
16 October 2018
Volume
12 - 2018
Edited by
Stephane Perrey, Université de Montpellier, France
Reviewed by
Bernd Figner, Radboud University Nijmegen, Netherlands; Julien Dubois, Cedars-Sinai Medical Center, United States; Frieder Michel Paulus, Universität zu Lübeck, Germany
Updates

Check for updates
Copyright
© 2018 de Haas.
This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.
*Correspondence: Benjamin de Haas, benjamindehaas@gmail.com
Disclaimer
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.