Manuscript #12553

Published on


Metadata

eLife Assessment

This manuscript reports an important new statistical method for calculating the significance of correlations between two time-series, which provides more accuracy than other methods when the data has few replicates. The proposed method solves a real-life problem that is frequently encountered and is broadly applicable to many realistic datasets in many experimental contexts. The technique is supported with compelling mathematical derivations as well as analysis of both computer-generated and previously published experimental data.

Reviewer #1 (Public review):

[Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review: Definitions and terminology have been made more precise. Additional analysis confirms the conclusions previously stated and clarifies concerns about the computational tractability of the method.]

Summary:

The manuscript puts forward a statistical method to more accurately report the significance of correlations within data. The motivation for this study is two-fold. First, the publication of biological studies demands the report of p-values, and it is widely accepted that p-values below the arbitrary threshold of 0.05 give the authors of such studies justification to draw conclusions about their data. Second, many biological studies are limited by the number of replicate samples that are feasible, with replicates of less than 5 typical. The authors report a statistical tool that uses a permute-match approach to calculate p-values. Notably, the proposed method reduces p-values from around 0.2 to 0.04 as compared to a standard permutation test with a small sample size. The approach is clearly explained, including detailed mathematical explanations and derivations. The advantage of the approach is also demonstrated through analysis of computer-generated synthetic data with specified correlation and analysis of previously published data related to fish schooling. The authors make a clear case that this method is an improvement over the more standard approach currently used and also demonstrate the impact of this methodology on the ability to obtain p-values that are the standard for biological research. Overall, this paper is very strong. While the subject matter seems somewhat specialized, I would make the case that this will be an important study that has broad general interest to readers. The findings are very general and applicable to many research contexts. Experimentalists also want to report accurate p-values in their work and better understand how these values are calculated. Although I believe the previous statement is true, I am not sure that many research groups doing biological work are reading specialized statistics journals regularly. Therefore, a useful and broadly applicable statistical tool is well placed in this journal.

Strengths:

The proposed method is broadly applicable to many realistic datasets in many experimental contexts.

The power of this method was demonstrated with both real experimental data and "synthetic" data. The advantages of the tool are clearly reported. The zebrafish data is a great example dataset.

The method solves a real-life problem that is frequently encountered by many experimental groups in the biological sciences.

The writing of the paper is surprisingly clear, given the technical nature of the subject matter. I would not at all consider myself a statistician or mathematician, but I found the text easy to follow. The authors did an impressive job guiding the reader through material that would often be difficult to grasp. The introduction was also well-written and clearly motivated the goals of the study.

Reviewer #2 (Public review):

Summary:

This paper presented a hypothesis testing procedure for the independence of two time-series that was potentially suitable for nonlinear dependence and for small-sample cases. This should bring potential benefits for biology data.

Strengths:

The test offers good flexibility for different kinds of dependence (through adjusting \rho) and seems to have good finite sample performance compared to the literature. The justification regarding the validity of the test procedure is clear.

Author response:

 

Public Reviews:

 

Reviewer #1 (Public review):

 

Summary:

 

The manuscript puts forward a statistical method to more accurately report the significance of correlations within data. The motivation for this study is two-fold. First, the publication of biological studies demands the report of p-values, and it is widely accepted that p-values below the arbitrary threshold of 0.05 give the authors of such studies justification to draw conclusions about their data. Second, many biological studies are limited by the number of replicate samples that are feasible, with replicates of less than 5 typical. The authors report a statistical tool that uses a permute-match approach to calculate p-values. Notably, the proposed method reduces p-values from around 0.2 to 0.04 as compared to a standard permutation test with a small sample size. The approach is clearly explained, including detailed mathematical explanations and derivations. The advantage of the approach is also demonstrated through analysis of computer-generated synthetic data with specified correlation and analysis of previously published data related to fish schooling. The authors make a clear case that this method is an improvement over the more standard approach currently used, and also demonstrate the impact of this methodology on the ability to obtain p-values that are the standard for biological research. Overall, this paper is very strong. While the subject matter seems somewhat specialized, I would make the case that this will be an important study that has broad general interest to readers. The findings are very general and applicable to many research contexts. Experimentalists also want to report accurate p-values in their work and better understand how these values are calculated. Although I believe the previous statement is true, I am not sure that many research groups doing biological work are reading specialized statistics journals regularly. Therefore a useful and broadly applicable statistical tool is well placed in this journal.

 

Strengths:

 

The proposed method is broadly applicable to many realistic datasets in many experimental contexts.

 

The power of this method was demonstrated with both real experimental data and "synthetic" data. The advantages of the tool are clearly reported. The zebrafish data is a great example dataset.

 

The method solves a real-life problem that is frequently encountered by many experimental groups in the biological sciences.

 

The writing of the paper is surprisingly clear, given the technical nature of the subject matter. I would not at all consider myself a statistician or mathematician, but I found the text easy to follow. The authors did an impressive job guiding the reader through material that would often be difficult to grasp. The introduction was also well-written and clearly motivated the goals of the study.

 

We appreciate the reviewer’s summary of our study and its strengths.

 

Weaknesses:

 

A few changes could be made if the manuscript is revised. I would consider all of these points minor, but the paper could be improved if these points were addressed.

 

(1) The caption of Figure 2 doesn't seem to mention panel D. Figure A-2 also does not mention C in the caption.

 

We apologize for this error, and thank you for catching it! The figure legends had missing or incorrect panel labels. This error has been corrected.

 

(2) Figure 2D is a little hard to follow. First, the definition of "Power" is not clear, and I couldn't find the precise definition in the text. Second, the legend for the different lines in 2D is only given in Figure A-2. Perhaps a portion of the caption for Figure 2 is missing?

 

We have added a definition of power in the main text:

 

“Although the permutation test, simultaneous permute-match test, and sequential permute-match test are all valid, they vary in power – the probability of detecting true dependence.”

 

We have clarified the use of “power” in legend of Fig 2 and clarified that the color key for Fig 2D is in Fig 2A. The relevant excerpt of the Fig 2 legend is copied here:

 

“(D) Statistical power for the permutation test and various permute-match tests as a function of the replicate number <i>n</i>, significance level <i>α</i>, and strength of dependence <i>r<sub>X, Y</sub></i>. Power was estimated as the proportion of simulations in which dependence was detected, calculated from 5000 simulations at each value of <i>r<sub>X, Y</sub></i> between <i>r<sub>X, Y</sub></i> = 0 and 0.54 in steps of size 0.01. At <i>r<sub>X, Y</sub></i> = 0, there is no dependence, so the curve at that point indicates the false positive rate rather than power. We chose the Pearson correlation coefficient as our correlation function <i>ρ</i>. See (A) for the color legend.”

 

We have also added dotted lines connecting the legend in panel A to the curves in panel D.

 

(3) The concept of circular variance for the fish data was heard to understand/visualize. The equation on line 326 did not help much. If there is a very simple picture that could be added near line 326 that helps to explain Ct and theta, that could be a big help for some readers who do not work on related systems. The analysis performed is understandable, the reader just has to accept that circular variance captions the degree of alignment of the fish.

 

We have replaced references to circular concentration with “mean resultant length”, which is the standard jargon for this term in circular statistics, and we have added an illustration.

 

(4) For the data discussed in Figure 3, I wasn’t 100% sure how the time windows were selected. In the caption, it says “time series to different lengths starting from the first frame”. So the 20 s time window was from t=0 to t= 20 s. Would a different result be obtained if a different 20 s window was chosen (from t = 4 min to t = 4 min 20 s just to give a specific example). I suppose by chance one of the time windows would give a pvalue less than the target 0.05, that wouldn’t be surprising. Maybe a random time window should be selected (although I am not indicating what was reported was incorrect)? A little more discussion on this aspect of the study may be helpful.

 

As suggested by the reviewer, we have redone the analysis of Figure 3D with random segments. This provides a more complete picture of how the chance of detecting a significant correlation varies with segment length. The main conclusion is unchanged: Perfect match tests reliably detect dependence across a wider range of segment lengths than the naive parametric alternative.

 

The relevant panel and an excerpt from the legend text are copied below.

 

“(D) Permute-match tests detected a significant correlation between speed and alignment more consistently than the parametric test. For a grid of lengths between 20 and 600 seconds we sampled 500 random segments of each length, each drawn from the first 600 seconds, and determined for each segment whether the parametric test and/or the two possible permute-match tests detected a significant (<i>p</i> ≤ 0.05) correlation. In the edge case of the maximum 600-second length, all 500 “random” segments were identical.”

 

Reviewer #2 (Public review):

 

Summary:

 

This paper presented a hypothesis testing procedure for the independence of two timeseries that was potentially suitable for nonlinear dependence and for small-sample cases. This should bring potential benefits for biology data.

 

Strengths:

 

The test offers good flexibility for different kinds of dependence (through adjusting \rho), and seems to have good finite sample performance compared to the literature. The justification regarding the validity of the test procedure is clear.

 

We appreciate the reviewer’s summary of key aspects of our manuscript.

 

Weaknesses:

 

(1) The size of the test is not guaranteed to (asymptotically) equal \alpha, which may damage the power.

 

We thank the reviewer for raising the issue of test size and power. We agree that a conservative test (one whose size can fall below alpha) may sacrifice power.

 

Our objective is distribution-free false-positive rate (FPR) control. That is, we wish to keep the FPR at or below alpha for every distribution of X and Y, because in our regime (nonstationary time series with few independent replicates) the scientist often cannot verify distributional assumptions. Inspired by the reviewer’s comment, we now show (new Proposition 14) that the perfect match probability can be made arbitrarily close to 1/<i>n</i><sup>!</sup>. As a consequence, any reported perfect match p-value below 1/<i>n</i><sup>!</sup> would break the distribution-free validity of the test.

 

A test that exploits distributional structure could likely access lower p-values; we have now explored how the empirical FPR of the permute-match test varies with the data-generating process (see our response to reviewer 2's recommendation 1 below).

 

(2) The computational time can be an issue for a moderately large sample size when calculating the X / Y-perfect match. It will be beneficial to include discussions on the implementations of the test.

 

We agree this is an important consideration. We have added the following text to the Discussion:

 

“The test appears computationally tractable for relevant sample sizes: Our implementation of the permute match<a href="https://cdn.elifesciences.org/public-review-media/103703/v2/Author-response-image-1.jpg"><img src="https://cdn.elifesciences.org/public-review-media/103703/v2/Author-response-image-1.jpg"></a> procedure completed a single test of dependence in the setting of Fig 2 with an average runtime of 3 seconds when n = 10 on a 2023 14-inch MacBook Pro with an M2 Pro processor and 16 GB RAM (see Source data 1). For <i>n</i> > 10, a standard permutation test already can report a <i>p</i>-value below 3 × 10<sup>−8</sup> so the perfect match test is likely unnecessary for typical applications.”

 

Recommendations for the authors:

 

Reviewer #1 (Recommendations for the authors):

 

A few more minor notes/comments:

 

As a personal preference, I like it when figures printed in grayscale retain their meaning (when possible). Just FYI, Figure 3D in grayscale is uninterpretable. Not saying a change is needed, just pointing it out.

 

We changed Fig 3D to address a comment above, and think the new version better distinguishes between the permute-match and permutation test results in greyscale.

 

On line 304, is that a lower bound or an upper bound? Maybe the issue is the probability mentioned on line 305 is not clear.

 

As this point is not the main focus of the investigation, we have rephrased it to make it less technical and eliminate the issue of which bound is in question.

 

“It seems likely that the tests could be further modified to report an even lower <i>p</i>-value when an <i>X</i>- and <i>Y</i> -perfect match occur simultaneously, as in Fig 3B. However, we have not investigated further and this problem is left for future efforts.”

 

I am not sure if this is a weakness, but the p-value changing depending on the choice of whether to apply the X-perfect match or Y-perfect match test first is fascinating. The authors did discuss this very issue at several points in the manuscript. It is slightly unsettling to me that there isn't an exact p-value for a given set of data. This one point gives me a new perspective on statistics.

 

In full transparency, I don't believe I have to background to thoroughly review the appendix. I did read through it and did not notice any errors, but I couldn't confidently say there are not any small mathematical errors or any logical flaws in the proofs. Some sections were not easy to follow (my own shortcomings, the writing appeared sufficient for more of an expert to understand).

 

We greatly appreciate the reviewer’s time and effort tackling an appendix outside their comfort zone.

 

Reviewer #2 (Recommendations for the authors):

 

(1) In the numerical experiment session, the authors should include the null situation, i.e., the performance of the test when X and Y are independent. This helps assess the size of the test.

 

We have added a section on size to our results section, copied below:

 

“The permute-match test’s false positive rate depends on the process tested. The permute-match test is conservative – meaning that its false positive rate can fall below the significance level – because both the permutation test and perfect match test are conservative. As discussed elsewhere [29], the permutation test is conservative when <i>α</i> is not one of its possible <i>p</i>-values and when ties may occur between the original correlation and shuffled correlations. Checking for a perfect match is similarly conservative. The actual probability of a false-alarm perfect match event can vary depending on the process being tested. To see this consider the permute-match<a href="https://cdn.elifesciences.org/public-review-media/103703/v2/Author-response-image-1.jpg"><img src="https://cdn.elifesciences.org/public-review-media/103703/v2/Author-response-image-1.jpg"></a> test in the setting where <i>n</i> = 3, where <i>α</i> = 0.05, and where <i>r<sub>X,Y</sub></i>= 0 (independent <i>X</i> and <i>Y</i>). Note that in this case, obtaining a <i>Y</i> -perfect match (and thus <i>p</i> = 1/n<sup>n</sup>) is necessary and sufficient to detect dependence since α is too low for detection by either the permutation test or the <i>p</i> = 2/n<sup>n</sup> leg of the permute-match test. In the linear system of Fig 2, we observed among 5000 simulations a detection rate of 0.0148, significantly below the upper bound of 1/3<sup>3</sup> (Figure 2 - Source data 1; one-tailed exact binomial test, <i>p</i> < 10<sup>−20</sup>). Conversely, in the nonlinear system of Fig S2, this same event (<i>Y</i> -perfect match under <i>n</i> = 3 and <i>r<sub>X,Y</sub></i> = 0) occurs with a detection rate of 0.0328, not significantly 9 below the upper bound of 1/3<sup>3</sup> (Figure S2 - Source data 1; one-tailed exact binomial test, <i>p</i> = 0.059). Thus, depending on the underlying process studied, the actual chance of a perfect match happening under independence may be near or significantly below the theoretical upper bound.”

 

(2) Some insights regarding the choice of rho should be provided. Especially, are there any examples that the classical test, such as the Pearson correlation or Granger causality test does not work?

 

We have redone the example of Appendix 4 with Pearson correlation, showing that Pearson correlation has substantially lower power than cross-map skill in this case (compare figures S2 and S3).

 

(3) Line 34 - 35, page 2: Correlations and causality should be separately considered. This sentence talks more about causality rather than correlation.

 

We appreciate the reviewer’s perspective and agree that correlation and causality are distinct.

 

We feel that pointing out the issue of spurious correlations is helpful to orient our readers, especially those from a broad scientific audience. In the text, we define “correlation” as a descriptive statistic (rather than normalized covariance), and later distinguish it from “dependence”, which has causal implications due to Reichenbach’s common cause principle. We believe this distinction provides a useful backdrop for practitioners who use statistical methods but are perhaps new to thinking deeply about statistical dependence.

 

(4) Please add some discussions on the situation that X_i depends on Y_{i - j} for some j > 0, which is associated with the setting of Granger causality test.

 

We have added the following to the discussion:

 

“No distributional assumptions are required, and the correlation function ρ can be completely arbitrary. For instance, <i>ρ</i> could include a lag to detect delayed dependence, or even evaluate the correlation strength at several lags and report the strongest among them [39].”