Analog similarity: best environmental analogs within a geographic envelope
Source:R/analog_similarity.R
analog_similarity.RdFinds, for each focal location, the environmental–nearest neighbor(s) in a
reference dataset that satisfy a specified geographic distance threshold.
This function is a wrapper that calls analog_search() using select = "knn_env".
Usage
analog_similarity(
x,
pool,
x_cov = NULL,
y = NULL,
weight = NULL,
coord_type = "auto",
geog,
env = NULL,
k = 20,
env_res_adj = "auto",
geog_res_adj = "auto",
cell_area_weight = "auto",
n_threads = NULL,
downsample = 1,
seed = NULL,
progress = FALSE
)Arguments
- x
Focal locations for which analogs will be found. Should be a matrix/data.frame with columns x, y, and environmental variables, or a SpatRaster with environmental variable layers.
- pool
The reference dataset to search for analogs. Either:
Matrix/data.frame with columns x, y, and environmental variables, or SpatRaster with environmental variable layers, OR
An
analog_indexobject created bybuild_analog_index()(for repeated queries).
- x_cov
Optional focal-specific covariance matrices for Mahalanobis distance calculations. Should be a matrix or data.frame with one row per focal location and one column per unique covariance component, or a SpatRaster with a layer for each component. For n environmental variables, there are n*(n+1)/2 unique components, ordered as: variances first (diagonals), then covariances (upper triangle by row).
- y
Optional vector, factor, matrix/data.frame, or SpatRaster giving values for each reference location (must have same number of rows/cells as
pool). Required for stats"sum","mean","weighted_sum","weighted_mean","regression", and"tabulate". Numeric for continuous stats; factor or coercible-to-factor (character, integer, logical) forstat = "tabulate".- weight
Optional pool site weights for use in aggregation. Numeric vector, single-column matrix/data.frame, or single-layer SpatRaster, with one value per row/cell of
pool. For aggregation stats like"weighted_mean","regression", etc., weights multiply through the weighted aggregation alongside any kernel weighting and cell-area weighting; they do not influence which analogs are selected byknn_*modes (selection remains distance-only). They are reported in pair mode as auser_weightcolumn. Values must be non-negative;NAis allowed and treated as 0 (the point is excluded from aggregation). DefaultNULLmeans no user-supplied weights.If you want to exclude a static subset of pool sites entirely, masking
pool(and any associatedy/covariates) upfront is more efficient than passingweight = 0for those sites, since the lattice index will not have to scan or distance-compute against them. Useweight = 0for cases where the mask varies per query against a shared index, or where some sites have a continuous weight and others should be excluded.- coord_type
Coordinate system type:
"auto"(default): Automatically detect from coordinate ranges."lonlat": Unprojected lon/lat coordinates (uses great-circle distance; assumesmax_geogis in km)."projected": Projected XY coordinates (uses planar distance; assumesmax_geogis in projection units).
- env, geog
Per-family distance treatment, each a
kernel()object (orNULL). A kernel bundles the hard distance threshold, the weighting kernel shape, and the kernel's scale for one family: environmental (env) or geography (geog).kernel(weight, theta, max, min)where:max: hard upper distance threshold — candidates beyond it (in that family's distance) are excluded. Forenv,maxmay be a single Euclidean radius or a per-variable vector of absolute-difference thresholds (length equal to the number of environmental variables); scalar environmental thresholds are in Mahalanobis units whenx_covis supplied. Forgeog,maxis a single radius (kilometers whencoord_type = "lonlat", projected units otherwise).min: hard lower distance threshold — candidates closer than it are excluded, so retained candidates form an annulusmin <= d <= max. Supported only forgeog(a single radius, same units as the geographicmax); settingminonenvis an error. Mainly used to impose a spatial buffer for cross-validation (seeanalog_cv()).weight: kernel shape for weighted aggregations —"uniform"(no distance weighting),"gaussian"(exp(-d^2 / (2 theta^2))), or"inverse"(1 / (1 + d / theta)). The overall kernel weight is the product of the two families' weights, so shapes may be mixed (e.g. an inverse environmental kernel with a Gaussian geographic kernel).theta: the kernel's scale (Gaussian bandwidth, or inverse half-weight distance). Seekernel_params()for calibrated values.
A
NULLkernel (the default for both) applies no threshold and no weighting for that family. Seekernel()for details.- k
Number of nearest analogs to return per focal location for kNN selection modes. Required when
selectis"knn_geog"or"knn_env"; must beNULLforselect = "all".- env_res_adj, geog_res_adj
Control the lattice search-index resolution of the environmental and geographic families, each a multiplier on a data-dependent default (targeting ~50 pool points per occupied bin, split between families by effective dimensionality, so it scales with pool size). Each is either:
A non-negative number:
1uses the default for that family, larger values are finer, smaller are coarser, and0deactivates it."auto"(the default for both): tune a single overall resolution scale by optimizing compute time on a subsample of focal points. If focal has relatively few rows, tuning is skipped. Not supported whendownsample < 1(set explicit numeric values instead).
A family that the query does not constrain (no corresponding
max_*and not the knn sort key) is automatically deactivated, overriding any explicit value (with a message), since binning an unconstrained family only costs time. Ignored ifpoolis ananalog_index(uses the index's resolution).- cell_area_weight
Controls cell-area weighting when
poolis a raster. One of"auto"(default; on for raster pools, off otherwise),TRUE(force on; errors ifpoolis not a SpatRaster), orFALSE(force off). Cell-area weights correct aggregation statistics for non-uniform cell areas (e.g. lonlat grids near the poles, or projected grids on non-equal-area projections); they are computed viaterra::cellSize()and normalized to mean 1. Whenpoolis a pre-builtanalog_index, this argument must agree with the index's stored configuration:cell_area_weight = FALSEerrors if the index was built with cell-area weighting on (rebuild the index instead).- n_threads
Optional integer number of threads to use for the computation. If
NULL(default), the global RcppParallel setting is used (seeRcppParallel::setThreadOptions).- downsample
Optional downsampling rate (0-1) for the reference pool, indicating the proportion of points to retain. Values < 1 reduce memory and improve speed at some cost to precision. Default is 1.0 (no downsampling). Ignored if
poolis a pre-built index. Whendownsample < 1, resolution must be set explicitly viageog_res_adj/env_res_adj(auto-tuning is not supported in this case; see those parameters for details).- seed
Optional random seed for reproducible downsampling. If
NULL(default), uses current R random state. Ignored ifpoolis a pre-built index ordownsample = 1.- progress
Logical; if
TRUE, display a progress bar during computation. Progress tracking works by splitting the focal dataset into chunks and processing them sequentially. Useful for large datasets. Default isFALSE.
Value
A data.frame, or a SpatRaster when x is one and k = 1.
Contains one row per focal-analog pair with index, x, y,
analog_index, analog_x, analog_y, env_dist, and
geog_dist. See analog_search() for full column conventions
and metadata() for attached metadata attributes.
Details
For each focal location, analog_similarity():
Identifies all reference points within
geog's max (km) (and optional environmental filter).Selects the
kclosest in environmental distance.
This is the natural "inverse" of analog_velocity: instead of finding
where the focal environmental moves geographically, it finds the closest environmentally
similar conditions that are geographically reachable.
Among other uses, this operation is often the first step in a traditional
analog impact modeling (AIM) analysis – though see analog_impact() for a
more complete AIM implementation.
See also
analog_search() for the underlying flexible analog search function;
tiled_analog_search() for memory-safe searches on large raster datasets.
Examples
if (FALSE) { # \dontrun{
# One-shot query
im <- analog_similarity(
x = clim$clim1,
pool = clim$clim2,
geog = kernel(max = 100),
k = 20
)
# With pre-built index (for repeated queries)
index <- build_analog_index(clim$clim2)
i1 <- analog_similarity(x = sites1, pool = index, geog = kernel(max = 100), k = 20)
i2 <- analog_similarity(x = sites2, pool = index, geog = kernel(max = 50), k = 10)
} # }