---
title: "Introduction to _Data over Space and Time_"
author: Cosma Shalizi
date: "25 August 2026, 36-740/620"
bibliography: locusts.bib
output:
  slidy_presentation:
    math_method:
      engine: mathjax
      url: https://bactra.org/mathjax/tex-svg.js
---

\[
\newcommand{\Expect}[1]{\mathbb{E}\left[ #1 \right]}
\newcommand{\Var}[1]{\mathrm{Var}\left[ #1 \right]}
\newcommand{\Cov}[1]{\mathrm{Cov}\left[ #1 \right]}
\]

```{r, include=FALSE}
# General set-up options
library(knitr)
opts_chunk$set(size="small",background="white", highlight=FALSE,
               cache=TRUE, autodep=TRUE,
               tidy=TRUE, warning=FALSE, message=FALSE,
               echo=FALSE)
```

## Data over Space and Time

- Statistical study of processes that unfold over space, or time, or both




## Data over Space and Time

- Statistical study of processes that unfold over space, or time, or both
    + Data over _space_:

```{r, fig.retina=NULL, out.width=700, echo=FALSE}
knitr::include_graphics("cdc-measles-space-2026.png")
```


(From [https://www.cdc.gov/measles/data-research/index.html] on 2026-08-25)


## Data over Space and Time

- Statistical study of processes that unfold over space, or time, or both
    + Data over space
    + Data over _time_:

```{r, fig.retina=NULL, out.width=700, echo=FALSE}
knitr::include_graphics("cdc-measles-time-2026-08-25.png")
```

(From [https://www.cdc.gov/measles/data-research/index.html] on 2026-08-25)


## Data over Space and Time

- Statistical study of processes that unfold over space, or time, or both
    + Data over space
    + Data over _time_:

```{r, fig.retina=NULL, out.width=700, echo=FALSE}
knitr::include_graphics("cdc-measles-time-since-1985.png")
```

(From [https://www.cdc.gov/measles/data-research/index.html] on 2026-08-25)


## Data over Space and Time

- Statistical study of processes that unfold over space, or time, or both
    + Data over space
    + Data over time
    + Data over space _and_ time:


```{r, fig.retina=NULL, out.width=300, echo=FALSE, fig.keep="all"}
par(mfrow=c(1,3))
knitr::include_graphics("cdc-measles-space-2024.png")
knitr::include_graphics("cdc-measles-space-2025.png")
knitr::include_graphics("cdc-measles-space-2026.png")
```


(From [https://www.cdc.gov/measles/data-research/index.html] on 2026-08-25)



## Data over Space and Time

- Statistical study of processes that unfold over space, or time, or both

> - Applications: History, science, and lots of practical policy & technology
> - Not everything:
>    + Good experiments
>     + Good surveys
>     + Wishful thinking

## Special Statistical Issues

> - The goal: learn how _here_ and _now_ relates to _there_ and _then_
> - The two problems:
>     1. Everything depends on everything else $\Rightarrow$ basic theory doesn't apply
>     2. We don't get multiple samples of time or of space
>     + $\Rightarrow$ $n=1$, always
> - Nonetheless, we will see how to do statistical inference with dependent, spatio-temporal data

## Course Mechanics {.smaller}

- Syllabus: [http://www.stat.cmu.edu/~cshalizi/dst/26]
    + All assignments will be posted there
    + Information about readings and other course resources ditto
- Canvas: turn in assignments, gradebook, some readings that can't go on the public web
- Piazza for question-answering


#### Office hours

- TBD: please fill out the poll you will get in e-mail by the end of the week


## Lectures

- Clarification, amplification, examples, alternatives
- _Complements_ to the readings
    + $\therefore$ Do the readings ahead of time
- Please ask questions
- I will _often_ ask you to solve some problem during lecture
  + not graded, _but_
  + students find the discussion of the solutions very helpful
  + and more helpful if you've tried it yourself
- _Not_ recorded


## Assignments/Grading

1. Weekly homework: 60%
2. Oral final exam: 20%
3. Scribing: 20% for 620, 10% for 740
4. Developing a research question: 0% for 620, 10% for 740


## Homework

- 60% of your total grade
- Data analysis & computing & a little theory
- Always turned in electronically via Canvas/Gradescope
- Due at 6 pm on Thursday every week
    + Except this first week (no homework)
    + and the last week (exams, no homework)
- Lowest 2 grades changed to 80/100
    + Unless that would lower your average
	+ This includes 0s for not turning in anything
- NO LATE HOMEWORK FOR ANY REASON
    + Turn in as many incomplete or rough versions as you like well ahead of the deadline
	
## Oral Final Exams

- 20% of your total grade
- One-on-one in my office during the last week of the mini
- 20 minutes or less
- I will start by asking you to describe the math of a method we have covered, _or_ how to apply it in a given situation, and go from there
- List of the 8--12 possible topics no less than 1 week before the exam (and probably much earlier)
- Open notes, open book, no computing devices
- Done during regular class time, with reserved 20 minute slots


## Scribing

- 20% of your grade for 620, 10% for 740
- Starting next week, every lecture gets a student scribe
- Take notes during lecture and turn them, within 1 week, into a readable PDF shared with the class
- You'll get my slides
- Please sign up on the spreadsheet you will get in e-mail

## Research question notes

- 10% for 740 (not assigned to 620)
- Develop a research _question_, and explain it in a short document of 1--5 pages
- Q can concern our methods, or their application
- You are not expected to do the research, just practice articulating a research question
- First draft mid-September, final draft end of mini, feedback in between


## Any questions?




## What are the big issues?

- We see $X$ at time $t$ and point $r$, $X(r,t)$, and what to know how it
relates to $Y$ at time $s$ and point $q$, $Y(q,s)$


Problems:

0. Basic statistical theory is about independent, identically distributed (IID) data.
1. But we (usually) only see _one_ realization of a whole process.
2. Every observation is dependent on every other observation.
3. Basic statistical theory says that $n=1$ and refuses to draw any inferences.

## How are we going to deal with these issues?

- Methods for describing relationships and finding patterns in the data
- Especially methods for _predicting_ $X(r,t)$ from $Y(q,s)$
- _Incorporate_ dependence into statistical theory, so we can say when methods will work
- Quantifying uncertain through modeling and simulation




## Why is this worth knowing?

- Italo Calvino, "All at One Point" (1963) imagines the world un-extended in time or space[^calvino]

[^calvino]: Calvino's _Cosmicomics_ is a precious part of our common cultural heritage.

- Otherwise, every branch of science deals with data spread over space and time
- Our examples will come from
    + Geology
    + Climatology
    + Meteorology
    + Ecology
    + Epidemiology
    + Demography
    + Economics
- Could (and might) add examples from neuroscience, physics, technology, policy, business, etc.



## So what are we going to cover? {.smaller}

- Smoothing and trends for spatio-temporal data
- Decomposing data into basis functions
- Linear prediction over space and time
- Inferring hidden states
- Quantifying uncertainty with cross-validation and bootstrap
- Organized by _method_, skip around on _subject matter_

#### What we will not cover

- ARIMA (etc.) models --- take 36-618
- Finance

### To get started...

- Something vivid, but more cheerful than measles


## Cherry blossoms in Kyoto

![](https://farm2.staticflickr.com/1726/27534132847_78448546b1.jpg)

> Cherries at the Hirano shrine in Kyoto [(David Montasco on flickr)](https://www.flickr.com/photos/david_montasco/27534132847/)

Flowering of cherry trees has been a central part of Japanese high art & culture for well over a millennium

## _Hanami_


```{r, fig.retina=NULL, out.width=400, echo=FALSE}
knitr::include_graphics("https://www.loc.gov/exhibits/cherry-blossoms/cultural-history/Assets/cb0019_enlarge.jpg")
```

> Kitao Shigemasa, _Sangatsu, Asukayam Hanami_ = _Third Lunar Month, Blossom Viewing at Asuka Hill_, c. 1776, via [Library of Congress](https://www.loc.gov/exhibits/cherry-blossoms/cherry-blossoms-in-japanese-cultural-history.html)

Notice the date in the title!

## This is data!

- Ancient diaries[^shonagon], poetry, etc., and modern newspapers, record _when_ cherry trees in Kyoto came into bloom
     + Kyoto because it's the ancient capital and has been a continuous seat of art & culture

[^shonagon]: The _Pillow Book of Sei Shonagon_ is a precious part of our common cultural heritage.


## Cherry blossoms track climate

- Japan gets cold

![](https://farm3.staticflickr.com/2682/32049377133_92a73aab96.jpg)

> Snow at the Hirano shrine [(yopparainokobito on flickr)](https://www.flickr.com/photos/kobikobi/32049377133/)

- Cherries only blossom when it gets warm enough

- $\therefore$ The date when cherry trees are in full flower tells us about how warm the year was
    + Date of first flowering is also informative, but less often recorded


## A data set

- Assembled by [Prof. Yasuyuki Aono](http://atmenv.envi.osakafu-u.ac.jp/aono/kyophenotemp4/)
    + Continued by Prof. [Genki Katata](https://cigs.canon/katata_date/260420sakura_katata.html)
    + Data points going back to the early 800s
    + Almost continuous for modern times
    + Re-formatted version at [http://www.stat.cmu.edu/~cshalizi/dst/26/data/kyoto-2026.csv]
- For each year, the day of the year (1--366) of full bloom
    + April 1 = `r 31+28+31+1` or `r 31+29+31+1`
    + Aono had to search out the ancient records, poems, histories, etc.
    + and convert dates before 1873 to our calendar[^calendar]


[^calendar]: The traditional Japanese calendar system (from 645) didn't have an accumulating count of years the way we do, but rather reckoned years by "name eras", so a given year would be called something like "year $k$ of the reign of Emperor So-and-so."  (Some emperors had more than one name era, and the name of the era was not the emperor's name but one he chose, but that was the basic idea.)  This is as though we called this year 2 of Trump II, called 2016 year 8 of Obama, etc.  Part of the work of compiling data like this is to keep track of when, in our terms, each name era began.  Japan in fact still has name eras for some official purposes (2026 is year 7 of the Reiwa era), but in 1873, as part of the Meiji Revolution, the government adopted the Gregorian calendar and the common era, the year-numbering scheme formerly known as AD/BC.  ("Meiji" is itself an era name.)  More-or-less similar schemes, where the count of years resets when the ruler changes, have been very common across the world; unending sequential year numbers seem to have been invented twice, by the Seleucid dynasty in what we now call the Middle East around 300 BCE, with [remarkable consequences](https://aeon.co/essays/when-time-became-regular-and-universal-it-changed-history), and in Central America, in the form of the "long count" calendar used by the Mayans and other civilizations and reckoning days since 11 August 3114 BCE.  (That calendar was certainly invented much more recently and we don't know why that was their zero-day.)


	

## A data set

```{r}
kyoto <- read.csv("http://www.stat.cmu.edu/~cshalizi/dst/26/data/kyoto-2026.csv")
plot(Flowering.DOY ~ Year.AD, data=kyoto, type="o",
     ylab="Day in year of full flowering", xlab="Year (AD)",
     main="Cherry blossoms in Kyoto",
     pch=16, cex=0.3)
```



## Problems {.smaller}

> - How to fill in the data _between_ the observations? (**interpolation**)
>     + When did the cherries bloom in 1015?
> - How to extend _beyond_ the observations? (**extrapolation**)
>     + Make a guess for 2030, _or_ for 800
> - How to remove measurement noise _in_ the observations? (**filtering**)
>     + Poets can be sloppy about dates
> - How to separate year-to-year **fluctuations** from longer-term **trends**?
>     + How do we model fluctuations?
> 	+ How do we model trends?
> - How seriously should we treat the warming since about 1800?  (**inference**)
>     + Could it be an artifact of denser observations?
>     + Does the climate just do stuff like this occasionally?
>     + How long do we need to wait to be confident?



## Smoothed cherry blossoms

```{r, echo=FALSE}
plot(Flowering.DOY ~ Year.AD, data=kyoto, type="l",
     ylab="Day in year of full flowering", xlab="Year (AD)",
     main="Cherry blossoms in Kyoto")
lines(with(na.omit(kyoto), smooth.spline(x=Year.AD, y=Flowering.DOY)), col="blue", lwd=4)
abline(lm(Flowering.DOY ~ Year.AD, data=kyoto), col="green", lwd=4)
```

Both curves come from averaging observed values



## More concrete problems

- How do we fit either the blue or the green curve to the data?
- How can we tell which curve is a better fit?
- There ought to be margins of error --- are the two curves even incompatible?
- What do we mean by a "trend" anyway?





## We're going to need to build some concepts

- Stochastic process = collection of random variables over time or space or
both, typically dependent, say $X(t)$
- Trend = central tendency of the process, say $\mu(t) = \mathbb{E}[X(t)]$
- Fluctuations = difference from the trend, $\epsilon(t) \equiv X(t) - \mu(t)$
- Fluctuations are tied to covariances:
\[
\Cov{X(t), X(s)} =\Expect{(X(t) - \mu(t)) (X(s)-\mu(s))} = \Expect{\epsilon(t) \epsilon(s)} = \Cov{\epsilon(t), \epsilon(s)}
\]
- How can we figure out the trend if we just see $X$ once?
- How can we figure out covariances if we just see $X$ once?



## By Thursday:

- Make sure you're on Canvas & Piazza for the course
- Sign up for scribing
- Fill out office hours poll 

(Let me know if you haven't gotten the invitations / links by 6 pm today)



## Take-aways

- Everything is statistically dependent on everything else
- Dependence means what happens _here_ and _now_ gives us information about _there_ and _then_, so we can predict
- Dependence means the statistical theory we've learned needs to be fixed
- For now, we will focus on _describing_ dependence
- Next time: Trends and smoothing


## Exercises (to think through, not hand in) {.smaller}

> A first warm-up on statistical theory with dependence

Throughout, suppose $\Expect{X(t)} = \mu$ for all $t$, and $\Var{X(t)} = \sigma^2 > 0$ for all $t$.  Define $A_n = n^{-1} \sum_{t=1}^{n}{X(t)}$.

1. (_Baby's first law of large numbers_) $\Cov{X(t), X(s)} = 0$ (unless $t=s$).  Explain why each step of this proof is justified:
\begin{eqnarray*}
\Expect{A_n} & = & \frac{n\mu}{n} = \mu\\
\Var{A_n} & = & \frac{1}{n^2}\sum_{t=1}^{n}{\Var{X(t)}} = \frac{n \sigma^2}{n^2} = \frac{\sigma^2}{n}\\
\Expect{(A_n - \mu)^2} & = & \frac{\sigma^2}{n}
\end{eqnarray*}
We have shown that $A_n \rightarrow \mu$; what is the sense of convergence?
2. Suppose $X(t) = \alpha + \beta X(t-1) + \epsilon(t)$, where $\Expect{\epsilon(t)} = 0$, $\Var{\epsilon(t)} = \rho^2$, and $\Cov{\epsilon(t), \epsilon(s)} = 0$ (unless $t=s$).  Prove the following:
    a. _If_ $\Expect{X(t)} = \mu$ for all $t$, _then_ $\mu=\frac{\alpha}{1-\beta}$.
    b. _If_ $\Var{X(t)} = \sigma^2$ for all $t$, _then_ $\sigma^2 = \frac{\rho^2}{1-\beta^2}$.  Why does this require $|\beta|<1$?
    c. Assuming (2.1) and (2.2), $\Cov{X(t), X(t+h)} = \sigma^2 \beta^h$.
3. (_Baby's first ergodic theorem_) Assume all the things from (2).  Show that $A_n \rightarrow \mu$ still, by showing that $\Expect{(A_n - \mu)^2} \rightarrow \frac{\sigma^2}{n\tau}$, and find an explicit expression for $\tau$
    + For purists: $\Expect{n(A_n-\mu)^2} \rightarrow \sigma^2/\tau$
