Scientists & Organizations

This technical note is published by FreeThink Technologies, Inc., the developer of ASAPprime® software.

This technical note describes the process of calculating error bars used in ASAPprime® (v6 and later). These error bars are the basis for the probability calculations.

Error of the Fit

When data are fit to a mathematical function, a linear or non-linear least squares regression analysis can be performed to minimize the distance between the best fit of the function and each of the values (the residual). In this process, the regression line (for linear functions) is the line that minimizes the sum of squared residuals. There are different ways to capture the error bars for that fit. In this Technical Note, the process used in ASAPprime® for taking the isoconversion times at each condition to determine the error bars for fitting the Arrhenius or modified Arrhenius equations is described. These error bars are then used to determine the corresponding probabilities based on that distribution.

For simplicity, the situation of temperature only will be considered here, i.e., no relative humidity term, since the expansion to add the relative humidity term should be straightforward to understand. Also, for simplicity we will discuss the growth of degradant but recognize that loss of potency will be handled in an analogous manner. This means that we are fitting a linear function of ln [(degspec−deginit)/tiso) (ln kiso) vs. 1/T, where deginit is the amount of degradant at time zero, degspec is the specification limit, tiso is the isoconversion time (time to fail) and T is the stress temperature in Kelvin. Fitting the line provides two parameters: ln A (the intercept) and Ea/R (the slope). For simplicity, we can consider the case of just three conditions and with the deginit = 0 and degspec = 0.5% (Table 1 and Figure 1).

Table 1. Example data set for error bar calculations.
Temp. (°C)1/Ttiso (days)ln kisoln kiso [best fit]Residual Squared
500.00309625−3.912−3.7390.030
600.0030034−2.079−2.4330.125
700.0029152−1.386−1.2030.034
Average0.003005−2.4590.189 (sum)
Arrhenius plot of ln k-iso versus 1/T for the three-condition example data set, showing the fitted regression line extrapolated from 70, 60 and 50 degrees Celsius down to 25 degrees Celsius.
Figure 1. Plot of example data set.

The fitted line represents the lowest sum of the square residuals for each of the three points. The mean square error (MSE) will equal the square root of the sum of the square residuals divided by the number of degrees of freedom. The square residuals are shown in Table 1. With three points and two parameters, there is only one degree of freedom. This means that the MSE equals 0.593/1 = 0.593.

To calculate the error bars for the fit, we need to calculate the standard error (SE) in the ln kiso. This is determined using the following equation:


Equation: SE equals the square root of 1 divided by (n minus 2), multiplied by the quantity — sum of (ln k-iso minus mean ln k-iso) squared, minus the square of the sum of (1/T minus mean 1/T) times (ln k-iso minus mean ln k-iso), all divided by the sum of (1/T minus mean 1/T) squared.

Here n is the number of conditions and the bars over a value indicate the average of that value. For the example in Table 1, SE = 0.437. This leads to the error bars for the isoconversion rates as shown in Figure 2.

Arrhenius plot of ln k-iso versus 1/T with dashed confidence bands above and below the fitted line, representing the error of the fit for the Table 1 example data.
Figure 2. Error bars representing the error of the fit calculation for the example data from Table 1.

From the error bars in the logarithm of the isoconversion rates, the degradation curve, with error bars, can be calculated at 25°C, as shown in Figure 3. Even though the fit to the data is relatively good (R2 = 0.94), there is still a wide range of possible degradant levels due to the small number of points used and the translation of logarithmic to normal rates.

Percent degradant versus time in years at 25 degrees Celsius, showing a blue mean prediction line with green upper and red lower lines for plus and minus one standard deviation based on the error of the fit.
Figure 3. Distribution of potential degradant levels as a function of time at 25°C based on the error of the fit for the example data set. The green and red lines represent plus/minus one standard deviation rates vs. the blue line (mean prediction).

Error Propagation from Isoconversion Times

Another situation can also occur, especially with only a small number of conditions. If the points lie directly on a line, there will be no error bar at all from the error of the fit. However, since each isoconversion time has an error bar, i.e., each point represents a part of a distribution, another type of error calculation becomes important to reflect the error bars of the isoconversion times themselves.

The error bars for isoconversion times reflect the error bars from the data (i.e., amount of degradation at a single time point and condition) coupled with the amount of extrapolation each requires to hit the specification limit. We can envision three scenarios: (1) The error bars for the isoconversion times are much larger than the error bars for the fit; (2) the error bars for the fit are much larger than the isoconversion time errors; and (3) both errors are comparable.

In calculating the error bars for degradation at lower temperatures propagating from the isoconversion time error bars at accelerated conditions, the fact that each isoconversion error bar is likely to be different makes a closed form solution problematic. Because of this, a Monte-Carlo simulation is carried out in ASAPprime® to determine this error bar. If we assume three conditions with isoconversion times corresponding to a perfect Arrhenius fit but now incorporate varying error bars at each condition (Table 2 and Figure 4), we can illustrate how this works.

Table 2. Example data set used to illustrate propagated calculation of error bars.
Temp. (°C)tiso (days)SD
50258
6042.5
7020.4
Arrhenius plot of ln k-iso versus 1/T in which the three data points sit exactly on the fitted line but carry horizontal error bars of differing width, largest at 60 degrees Celsius and smallest at 70 degrees Celsius.
Figure 4. Plot of example data set used to illustrate propagated calculation of error bars.

As can be seen, the errors of the points are much larger than any error from the fit itself. We can generate a random distribution of isoconversion times for each temperature with the indicated average and standard deviation. We can use a total of 50 simulations as illustrated in Table 3 (ASAPprime® defaults to 2500 simulations).

Table 3. Monte-Carlo simulation for example data from Table 2.
#50°C60°C70°C
120.83.41.8
232.90.32.6
322.76.42.5
47.58.01.9
519.67.12.1
632.24.62.1
724.46.12.1
835.03.62.5
918.91.21.4
1018.24.61.9
1127.57.12.1
1221.16.11.7
1323.83.81.9
1422.87.52.4
1513.81.71.6
1622.90.72.1
1729.18.52.3
1819.86.12.2
1915.95.62.3
2023.44.32.2
2132.95.71.8
2224.13.92.1
2318.38.22.4
2426.63.12.1
2532.48.32.5
2618.24.52.3
2716.72.72.5
2847.75.71.3
2931.01.92.1
3024.15.02.4
3124.50.91.8
3219.05.31.7
3316.32.12.2
3416.50.72.4
3546.01.41.9
3643.13.82.5
3716.55.32.0
3830.24.32.0
3925.98.61.8
4029.26.32.0
4134.25.11.9
4234.91.81.8
4320.90.81.6
4424.57.12.7
4526.01.21.8
4633.03.71.7
4721.85.22.6
4834.74.22.2
4934.13.01.6
5023.61.01.6
Avg.25.64.42.1
SD8.12.40.3

The 50 simulations have averages and standard deviations that are close to the targets from Table 2 and are expected to get even closer with greater numbers of simulations. For each of the 50 simulations, the isoconversion times can be fit to the Arrhenius equation, and the resulting parameters for ln A and Ea provide a distribution of values which can be approximated as a normal distribution. The results are shown in Table 4.

Table 4. Results of fifty-iteration Monte-Carlo simulation for data from Table 3.
Monte-Carlo MeanMonte-Carlo Standard DeviationRegression Line Mean
ln A39.16.340.5
Ea (kcal/mol)27.44.227.9

As can be seen, the regression means and Monte-Carlo means are not quite the same because of the limited number of simulations used in this example. In ASAPprime®, the regression mean is used to center the distribution with the standard deviations calculated from the Monte-Carlo simulation. The confidence bands can be seen in Figure 5 and can be compared to those in Figure 2 based on the error of the fit.

Arrhenius plot of ln k-iso versus 1/T with dashed confidence bands derived from Monte-Carlo propagation of the isoconversion rate error bars, widening toward the 25 degrees Celsius extrapolation.
Figure 5. Confidence bands for example data set based on propagating the error bars for the isoconversion rates using a Monte-Carlo calculation.

Using these confidence bands, the standard deviation for the isoconversion rates at 25°C can be used to generate the projected degradation error bars shown in Figure 6.

Percent degradant versus time in years at 25 degrees Celsius, with a blue mean prediction line, a steep green upper line and a near-flat red lower line, based on Monte-Carlo propagated error bars.
Figure 6. Distribution of potential degradant levels as a function of time at 25°C based on the Monte-Carlo propagated error bars for the data set. The green and red lines represent plus/minus one standard deviation rates from the blue line (mean prediction).

Combining Errors

In a sense, one would expect that if a fitting model (in this case an Arrhenius fit) is correct, the error bars for that model should be encompassed naturally within the error bars for the points used. In other words, if the fit to the model with its error bars does not encompass the points with their error bars, one would have to consider that the model is not correct for that data set or that the error bars used are too narrow. This was the assumption implicit with ASAPprime® versions earlier than v6. However, a bad fit to the model could easily result in a tight error bar if the isoconversion errors are small, giving many ASAPprime® users confident values that were not accurate. For example, if one condition showed a discontinuity due to melting, the data for each condition could be tight, yet the overall model fit would be poor. Removing the point would result in a more accurate model but not necessarily a smaller error bar. To remedy this situation, we combine both the error of the fit and the isoconversion error propagation into a new, effective error bar that takes both types of error bars into account. This is done by using the following calculation:


SDtotal = √SDfit2 + SDpropagated2

In the present example, the combination of error bars produces the behavior shown in Figure 7. At 25°C, this corresponds to the degradant growth curve shown in Figure 8. As can be seen, the Monte-Carlo propagated error is dominant in this case; however, the additional error due to the goodness of fit is noticeable in the error bars. For example, at three years, the one standard deviation higher than the mean predicted value is 1.05% using just the error of the fit, 1.56% using just the propagated error, and 1.71% using the square root of the sum of the square errors.

Arrhenius plot of ln k-iso versus 1/T with dashed confidence bands that combine the error of the fit and the Monte-Carlo propagated isoconversion error, wider than either source alone.
Figure 7. Confidence bands for example data set based on the square root of the sum of the square errors from the error of the fit and the propagated error bars from the isoconversion rates.
Percent degradant versus time in years at 25 degrees Celsius based on the combined error, with the green upper line reaching about 1.7 percent at three years against a blue mean prediction near 0.7 percent.
Figure 8. Distribution of potential degradant levels as a function of time at 25°C based on the square root of the sum of the square errors from the error of the fit and the Monte-Carlo propagated error bars for the data set. The green and red lines represent plus/minus one standard deviation rates from the blue line (mean prediction).

Summary

Error bar calculations need to reflect the uncertainty in modeling predictions. For ASAPprime®, the two major modeling errors are the error of the fit from the isoconversion rates (i.e., how well the data fit the model) and the error bar propagation from the isoconversion errors. The latter term reflects a combination of the experimental errors and the degree of extrapolation used to estimate the isoconversion rate at each condition. The software combines the two error sources using the square root of the sum of the squared errors. The overall error bars are then used to determine the probability of passing. The combination results in a more conservative estimate of shelf-life compared to using only the propagated error that was used in earlier versions of ASAPprime®.

Read the original technical note
Download the full ASAPprime® Software Technical Note as a PDF.
Download the PDF