We all wish we had a crystal ball. For an importer-exporter, knowing the value of the Indian Rupee against the US Dollar next month could mean the difference between profit and loss. For a business, forecasting sales can drive inventory, hiring, and strategy. While we don’t have magic, we do have powerful statistical tools. One of the most robust and widely respected methods for this kind of crystal-gazing is ARIMA modelling. Itโs a cornerstone of time series forecasting, and thanks to the R programming language, itโs more accessible than ever. This guide will walk you through the entire process, from data to forecast, using R as our workshop.
Table of Contents
- Getting to know our tools: R and the ‘forecast’ package
- The game plan: An ARIMA modelling workflow
- Step 1: Identification (What kind of model fits our data?)
- Step 2: Estimation (Building the model)
- Step 3: Diagnostic checking (Did we build a good model?)
- Let’s make it real: Modelling the INR-USD exchange rate
- Step 1 (Example): Identification
- The modern shortcut: auto.arima()
- Steps 2 & 3 (Example): Fitting and checking
- Diagnostic checking and forecasting in R (The payoff)
Getting to know our tools: R and the ‘forecast’ package
First, let’s meet our tools. R is a free, open-source software environment for statistical computing and graphics. It’s the language of choice for statisticians, data scientists, and academics because of its power, flexibility, and a massive community that builds and shares add-ons called “packages.”
Think of R as a brand-new, top-of-the-line smartphone. Itโs powerful on its own, but it truly shines when you install apps. In R, these “apps” are called packages. For our forecasting mission, the most essential package is ‘forecast’. Developed by one of the world’s leading forecasting experts, Rob Hyndman, this package simplifies and automates many of the most complex parts of time series analysis. It contains the functions that will do the heavy lifting for us, making an advanced technique like ARIMA feel much less intimidating.
We will also use the ‘tseries’ package, which contains a crucial test for one of our first steps. You can easily install them in R by running install.packages("forecast") and install.packages("tseries").
The game plan: An ARIMA modelling workflow
We can’t just feed raw data into a machine and ask for a future number. We need a process. The one we’ll follow is the celebrated Box-Jenkins method, named after the statisticians who developed it. Itโs an iterative, three-stage process that ensures we build a model that is both accurate and reliable. The three stages are: Identification, Estimation, and Diagnostic Checking.
Step 1: Identification (What kind of model fits our data?)
This is the most crucial (and most human) part of the process. We start by acting like a detective, looking for clues in our data.
Look at the data
Before you do anything else, plot your data. A simple line chart (using the plot() function in R) is the single most important diagnostic tool. We are looking for two main things:
- Trend: Is the data generally moving upwards or downwards over time?
- Seasonality: Are there repeating, predictable cycles (e.g., sales are always high in December, low in January)?
The concept of stationarity
ARIMA models have a very important requirement: they need the data to be stationary. A stationary time series is one whose statistical properties (like the mean and variance) are constant over time. Think of it this way: it’s incredibly difficult to forecast a river that is constantly getting wider and flowing faster. It’s much easier to predict a calm, steady canal. Our job is to turn our volatile river into a calm canal.
Most real-world economic data, like an exchange rate, is non-stationary. The INR-USD rate, for example, has a general upward trend over the last few decades. To fix this, we use a technique called differencing. This is the ‘I’ (Integrated) in ARIMA. Instead of modelling the price, we model the *change* in price from one day to the next. This new series (the day-to-day change) will likely be stationary, hovering around a mean of zero. The number of times we have to difference the data to make it stationary is our ‘d’ parameter.
While looking at a plot is a good start, we can formally test for stationarity using the Augmented Dickey-Fuller (ADF) test (from the ‘tseries’ package). In this test, a small p-value (e.g., < 0.05) gives us the confidence to say our data is stationary.
Finding ‘p’ and ‘q’ with ACF/PACF
Once our data is stationary (our ‘d’ is found), we need to find two more parameters. This is where ARIMA (AutoRegressive Integrated Moving Average) gets its other letters.
- AR (AutoRegressive) – ‘p’: This part of the model assumes that the current value is dependent on its own past values. The ‘p’ is the number of past (lagged) values it depends on.
- MA (Moving Average) – ‘q’: This part assumes that the current value is dependent on the “shocks” or forecast errors from the past. The ‘q’ is the number of past errors it depends on.
To find `p` and `q`, we look at two more plots:
- Autocorrelation Function (ACF): This plot shows the correlation of the series with its past values at different lags.
- Partial Autocorrelation Function (PACF): This plot is similar, but it “controls for” the correlations at shorter lags.
Reading these plots is a bit of an art, but the general rules of thumb are:
- An AR(p) model (purely autoregressive) will have an ACF that “trails off” (slowly dies down) and a PACF that “cuts off” (drops to zero) after lag `p`.
- An MA(q) model (purely moving average) will have an ACF that “cuts off” after lag `q` and a PACF that “trails off.”
- An ARMA(p,q) model (a mix) will often have both plots trail off.
This step can be tricky and subjective. But don’t worry, R has a powerful shortcut for us, which we’ll see in the example.
Step 2: Estimation (Building the model)
Once we have our three parameters (p, d, q), this step is the easiest. We simply pass them to R’s arima() function. For example, if we identified an ARIMA(2,1,2) model, we’d just run: fit <- arima(my_data, order = c(2, 1, 2)) R will then use maximum likelihood estimation (MLE) to find the best possible coefficients for our AR(2) and MA(2) terms. It's like telling the master craftsman, "I need a cabinet with two shelves and two drawers," and R goes and builds it perfectly.
Step 3: Diagnostic checking (Did we build a good model?)
We're not done yet. After building the model, we must check our work. How? By looking at the residuals. The residuals are the "leftovers"-the difference between our model's predictions and the actual historical data. A good model should capture all the patterns in the data, leaving behind only random, unpredictable noise. This iterative cycle of identification, estimation, and checking is the core of the Box-Jenkins methodology.
The residuals should look like white noise. Think of it as the static on an old TV-no patterns, no trends, just pure randomness. We test this using a formal test called the Ljung-Box test. In R, the command is Box.test(fit$residuals, type = "Ljung-Box"). For this test, we want a high p-value (e.g., > 0.05). A high p-value means we *cannot* reject the null hypothesis that the residuals are independent and uncorrelated. In simple terms: a high p-value is good news, it means our residuals are just random noise, and our model is a good fit.
Let's make it real: Modelling the INR-USD exchange rate
Let's walk through this process with our real-world example. We want to forecast the daily INR-USD exchange rate. This is a classic problem in economics, and ARIMA is a well-suited tool for it.
Step 1 (Example): Identification
First, we load our data (let's say it's in a time series object called inr_usd_ts) and plot it. plot(inr_usd_ts) We'd immediately see a non-stationary upward trend. So, we run an ADF test: adf.test(inr_usd_ts) The p-value would be high (e.g., 0.7), confirming non-stationarity. So, we take the first difference: diff_data <- diff(inr_usd_ts) We plot diff_data and find it now hovers around zero. The ADF test on diff_data gives a tiny p-value (e.g., 0.01). Perfect. Our data is now stationary, and we've found our first parameter: d = 1.
Now, we look at the ACF and PACF plots for diff_data. acf(diff_data) pacf(diff_data)
Let's say we see both plots trailing off, with significant spikes at lags 1 and 2 in the PACF and lags 1 and 2 in the ACF. This suggests a mixed ARMA model. Based on our prompt, we might hypothesize an ARIMA(2,1,2) model.
The modern shortcut: auto.arima()
Guessing p and q from plots can be difficult. This is where the 'forecast' package becomes our best friend. We can use one powerful function to do this entire identification step for us: auto.arima(). auto_fit <- auto.arima(inr_usd_ts)
This single command runs a sophisticated algorithm that tests different combinations of p, d, and q. It automatically performs unit root tests to find `d`, and then minimizes an information criterion (like AICc) to find the "best" and most parsimonious `p` and `q`. It's a fantastic and reliable way to get a great model to start with. For our example, let's say auto.arima() also suggests an ARIMA(2,1,2) model.
Steps 2 & 3 (Example): Fitting and checking
Now we have our model, we can use the Arima() function from the 'forecast' package (note the capital 'A', which is a more robust version of the base `arima()` function). fit <- Arima(inr_usd_ts, order = c(2, 1, 2))
Next, we check the residuals. The 'forecast' package gives us another great shortcut: checkresiduals(fit)
This function provides a plot of the residuals, their ACF (which should have no significant spikes), and automatically runs the Ljung-Box test. We look at the test output: Ljung-Box test data: Residuals from ARIMA(2,1,2) Q* = 20.5, df = 21, p-value = 0.49
The p-value is 0.49. This is much higher than 0.05, so we are thrilled! This confirms our residuals are white noise. Our model is good. If the p-value were low, we would need to go back, check our (p,d,q) orders, and try a different model.
Diagnostic checking and forecasting in R (The payoff)
With our model (fit) built and validated, we can finally do what we came here for: forecast the future. This is the easiest part of all. We use the forecast() function.
Let's say we want to forecast the exchange rate for the next 30 days. my_forecast <- forecast(fit, h = 30)
This my_forecast object contains everything we need: the point forecasts (the single best guess), and also the 80% and 95% confidence intervals. The best way to see this is to plot it. plot(my_forecast)
The plot will show the historical data, followed by our 30-day forecast. The forecast won't be a single line; it will be a "fan" or "cone" that gets wider as it goes further into the future. This is a crucial and honest feature. It visually represents our uncertainty. Forecasting one day ahead is much easier than forecasting 30 days ahead, and the widening confidence intervals reflect that.
This entire, powerful workflow, from plotting raw data to forecasting, is the complete Box-Jenkins cycle. By following these steps-Identify, Estimate, Check, and Forecast-we've moved from just looking at data to building a robust statistical model that can help us navigate the uncertainties of the future.
What do you think? How might sudden economic events, like an RBI policy announcement, affect the accuracy of a simple ARIMA forecast? Beyond exchange rates, what other financial or economic data would you be interested in modelling with this process?
References
- https://en.wikipedia.org/wiki/Box%E2%80%93Jenkins_method
- https://www.rdocumentation.org/packages/stats/versions/3.6.2/topics/Box.test
- https://www.ies.gov.in/pdfs/SeminarPaper-ArushiGupta.pdf
- https://otexts.com/fpp2/arima-r.html
- https://www.geeksforgeeks.org/r-language/time-series-analysis-using-arima-model-in-r-programming/
Leave a Reply