---
title: "splitGraph: From Metadata to Leakage-Aware Split Design"
author: "Selçuk Korkmaz"
date: "`r Sys.Date()`"
output:
  rmarkdown::html_vignette:
    toc: true
vignette: >
  %\VignetteIndexEntry{splitGraph: From Metadata to Leakage-Aware Split Design}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r setup, include=FALSE}
knitr::opts_chunk$set(
  collapse = TRUE,
  comment = "#>",
  message = FALSE,
  warning = FALSE,
  eval = TRUE
)

package_root <- if (file.exists("../DESCRIPTION")) ".." else "."
if (requireNamespace("pkgload", quietly = TRUE) &&
    file.exists(file.path(package_root, "DESCRIPTION"))) {
  pkgload::load_all(package_root, export_all = FALSE, helpers = FALSE, quiet = TRUE)
} else {
  library(splitGraph)
}

or_empty <- function(x) {
  if (is.null(x)) character() else x
}
```

## Why `splitGraph` exists

Leakage in biomedical evaluation workflows often comes from dataset structure rather
than from an obvious coding mistake. Two samples may look independent in a
model matrix while still sharing the same subject, batch, study, timepoint, or
feature provenance. If those relationships stay implicit, train/test
separation can look correct while violating the scientific separation you
actually intended.

`splitGraph` exists to make those relationships explicit before evaluation. It
turns metadata into a typed dependency graph that can be:

- validated for structural and leakage-relevant problems
- queried to inspect hidden overlap and provenance
- converted into deterministic split constraints
- translated into a stable, tool-agnostic split specification through the
  `split_spec` class and `as_split_spec()` / `validate_split_spec()` API
- serialized to a schema-versioned JSON that any language can consume (an R
  session, or the shipped Python reader driving scikit-learn)

The package is intentionally narrow. It does not fit models, run preprocessing
pipelines, or generate resamples by itself. Its job is to represent dependency
structure clearly enough that downstream evaluation can be trustworthy.

## A realistic toy dataset

The example below includes exactly the kinds of relationships that usually
matter for leakage-aware evaluation:

- repeated subjects (`P1` and `P2`)
- reused batch (`B1`)
- one subject (`P2`) appearing across studies
- explicit time ordering with one sample missing time metadata
- a feature set derived at the full-dataset scope

```{r metadata}
meta <- data.frame(
  sample_id = c("S1", "S2", "S3", "S4", "S5", "S6"),
  subject_id = c("P1", "P1", "P2", "P3", "P4", "P2"),
  batch_id = c("B1", "B2", "B1", "B3", NA, "B1"),
  study_id = c("ST1", "ST1", "ST1", "ST2", "ST3", "ST2"),
  timepoint_id = c("T0", "T1", "T0", "T2", NA, "T1"),
  assay_id = c("RNAseq", "RNAseq", "RNAseq", "RNAseq", "Proteomics", "RNAseq"),
  featureset_id = c("FS_GLOBAL", "FS_GLOBAL", "FS_GLOBAL", "FS_GLOBAL", "FS_PROT", "FS_GLOBAL"),
  outcome_id = c("O_case", "O_case", "O_ctrl", "O_case", "O_ctrl", "O_ctrl"),
  stringsAsFactors = FALSE
)

meta
```

This is still a small example, but it already contains enough structure to
make naive random splitting risky.

## Fast path: `graph_from_metadata()`

When your metadata already uses the canonical column names (`sample_id`,
`subject_id`, `batch_id`, `study_id`, `timepoint_id`, `time_index`,
`assay_id`, `featureset_id`, `site_id`, `region_id`, `platform_id`,
`outcome_id` / `outcome_value`), `graph_from_metadata()` does ingestion, typed
node construction, canonical edge construction, and optional
`timepoint_precedes` derivation in a single call. Any canonical column that is
absent is simply skipped, so you only supply the structure you have:

```{r fast-path}
quick_graph <- graph_from_metadata(
  data.frame(
    sample_id    = c("S1", "S2", "S3", "S4", "S5", "S6"),
    subject_id   = c("P1", "P1", "P2", "P2", "P3", "P3"),
    batch_id     = c("B1", "B2", "B1", "B2", "B1", "B2"),
    timepoint_id = c("T0", "T1", "T0", "T1", "T0", "T1"),
    time_index   = c(0, 1, 0, 1, 0, 1),
    outcome_id   = c("ctrl", "case", "ctrl", "case", "case", "ctrl")
  ),
  graph_name = "quick_demo"
)

quick_graph
```

If you pass `outcome_value` as a numeric column (e.g. `0`/`1`), the package
will create one `Outcome` node per distinct numeric value (`outcome:0`,
`outcome:1`) and emit a warning. For most workflows you want a character
class label, so prefer `outcome_id`.

The rest of this vignette uses the explicit constructor path because it lets
us show node attributes (`time_index`, `visit_label`, `platform`,
`derivation_scope`) and edges that `graph_from_metadata()` does not build
for you on its own — specifically `featureset_generated_from_study` and
`featureset_generated_from_batch`. (`subject_has_outcome` is the one
non-default outcome edge `graph_from_metadata()` will build for you when
you pass `outcome_scope = "subject"`.) Use `graph_from_metadata()` when
the canonical columns are enough; use the explicit path when you need
custom attributes or feature-provenance edges.

## Ingest metadata and build typed nodes and edges

The first step is to standardize metadata and then turn each entity type into
canonical graph nodes. Sample-level relations become typed edges.

```{r construction}
meta <- ingest_metadata(meta, dataset_name = "VignetteDemo")

sample_nodes <- create_nodes(meta, type = "Sample", id_col = "sample_id")
subject_nodes <- create_nodes(meta, type = "Subject", id_col = "subject_id")
batch_nodes <- create_nodes(meta, type = "Batch", id_col = "batch_id")
study_nodes <- create_nodes(meta, type = "Study", id_col = "study_id")

time_nodes <- create_nodes(
  data.frame(
    timepoint_id = c("T0", "T1", "T2"),
    time_index = c(0L, 1L, 2L),
    visit_label = c("baseline", "follow_up", "late_follow_up"),
    stringsAsFactors = FALSE
  ),
  type = "Timepoint",
  id_col = "timepoint_id",
  attr_cols = c("time_index", "visit_label")
)

assay_nodes <- create_nodes(
  data.frame(
    assay_id = c("RNAseq", "Proteomics"),
    modality = c("transcriptomics", "proteomics"),
    platform = c("NovaSeq", "Orbitrap"),
    stringsAsFactors = FALSE
  ),
  type = "Assay",
  id_col = "assay_id",
  attr_cols = c("modality", "platform")
)

featureset_nodes <- create_nodes(
  data.frame(
    featureset_id = c("FS_GLOBAL", "FS_PROT"),
    featureset_name = c("global_rna_signature", "proteomics_panel"),
    derivation_scope = c("per_dataset", "external"),
    feature_count = c(500L, 80L),
    stringsAsFactors = FALSE
  ),
  type = "FeatureSet",
  id_col = "featureset_id",
  attr_cols = c("featureset_name", "derivation_scope", "feature_count")
)

outcome_nodes <- create_nodes(
  data.frame(
    outcome_id = c("O_case", "O_ctrl"),
    outcome_name = c("response", "response"),
    outcome_type = c("binary", "binary"),
    observation_level = c("subject", "subject"),
    stringsAsFactors = FALSE
  ),
  type = "Outcome",
  id_col = "outcome_id",
  attr_cols = c("outcome_name", "outcome_type", "observation_level")
)

subject_edges <- create_edges(
  meta, "sample_id", "subject_id",
  "Sample", "Subject", "sample_belongs_to_subject"
)

batch_edges <- create_edges(
  meta, "sample_id", "batch_id",
  "Sample", "Batch", "sample_processed_in_batch",
  allow_missing = TRUE
)

study_edges <- create_edges(
  meta, "sample_id", "study_id",
  "Sample", "Study", "sample_from_study"
)

time_edges <- create_edges(
  meta, "sample_id", "timepoint_id",
  "Sample", "Timepoint", "sample_collected_at_timepoint",
  allow_missing = TRUE
)

assay_edges <- create_edges(
  meta, "sample_id", "assay_id",
  "Sample", "Assay", "sample_measured_by_assay"
)

featureset_edges <- create_edges(
  meta, "sample_id", "featureset_id",
  "Sample", "FeatureSet", "sample_uses_featureset"
)

outcome_edges <- create_edges(
  data.frame(
    subject_id = c("P1", "P2", "P3", "P4"),
    outcome_id = c("O_case", "O_ctrl", "O_case", "O_ctrl"),
    stringsAsFactors = FALSE
  ),
  "subject_id", "outcome_id",
  "Subject", "Outcome", "subject_has_outcome"
)

precedence_edges <- create_edges(
  data.frame(
    from_timepoint = c("T0", "T1"),
    to_timepoint = c("T1", "T2"),
    stringsAsFactors = FALSE
  ),
  "from_timepoint", "to_timepoint",
  "Timepoint", "Timepoint", "timepoint_precedes"
)

featureset_from_study <- create_edges(
  data.frame(
    featureset_id = "FS_GLOBAL",
    study_id = "ST1",
    stringsAsFactors = FALSE
  ),
  "featureset_id", "study_id",
  "FeatureSet", "Study", "featureset_generated_from_study"
)

featureset_from_batch <- create_edges(
  data.frame(
    featureset_id = "FS_GLOBAL",
    batch_id = "B1",
    stringsAsFactors = FALSE
  ),
  "featureset_id", "batch_id",
  "FeatureSet", "Batch", "featureset_generated_from_batch"
)
```

The node and edge tables are canonical and typed. The package assigns globally
unique node IDs such as `sample:S1` and `subject:P1`, so different entity
types cannot collide accidentally.

```{r construction-output}
sample_nodes
as.data.frame(sample_nodes)[, c("node_id", "node_type", "node_key", "label")]

edge_preview <- do.call(rbind, lapply(
  list(
    subject_edges, batch_edges, study_edges, time_edges,
    assay_edges, featureset_edges, outcome_edges,
    precedence_edges, featureset_from_study, featureset_from_batch
  ),
  as.data.frame
))

edge_preview[, c("from", "to", "edge_type")]
```

The node table shows the canonical sample IDs that everything else refers to.
The edge table shows the package's central design choice: dependency structure
is explicit, typed, and inspectable.

## Assemble the dependency graph

```{r graph}
graph <- build_dependency_graph(
  nodes = list(
    sample_nodes, subject_nodes, batch_nodes, study_nodes,
    time_nodes, assay_nodes, featureset_nodes, outcome_nodes
  ),
  edges = list(
    subject_edges, batch_edges, study_edges, time_edges,
    assay_edges, featureset_edges, outcome_edges,
    precedence_edges, featureset_from_study, featureset_from_batch
  ),
  graph_name = "vignette_graph",
  dataset_name = attr(meta, "dataset_name")
)

graph
summary(graph)
```

At this point the package has a single `dependency_graph` object with both
tabular and `igraph` representations behind it. The summary is useful because
it tells you exactly which entity types and relation types are present before
you derive any split rules.

Both representations stay reachable. `graph$nodes` and `graph$edges` are the
tables; `as_igraph()` hands you the `igraph` object, with `node_type` and
`node_key` on the vertices and `edge_type` on the edges, so anything igraph can
compute is available without leaving the typed model:

```{r as-igraph}
ig <- as_igraph(graph)
ig
igraph::vertex_attr_names(ig)
```

### Visualize the typed structure

`plot()` renders the graph with a typed, layered layout: `Sample` on top,
peer dependencies (`Subject`, `Batch`, `Study`, `Timepoint`) in the middle
band, `Assay`/`FeatureSet` below that, and `Outcome` at the bottom. Node
colors are keyed to type and an auto-generated legend is drawn by default.

```{r plot, fig.width = 7, fig.height = 5}
plot(graph)
```

Useful options:

```{r plot-options, eval = FALSE}
plot(graph, layout = "sugiyama")         # alternative hierarchical layout
plot(graph, show_labels = FALSE)         # hide labels on dense graphs
plot(graph, legend = FALSE)              # suppress the legend
plot(graph, legend_position = "bottomright")
plot(graph, node_colors = c(Sample = "#000000"))
```

Two other views answer different questions. `focus = "sample_projection"` drops
the dependency nodes and draws only the samples, joined whenever they share a
dependency of a type in `via`: this is a picture of the grouping a composite
derivation would produce, and a connected blob is a warning that the split has
little room left. `focus = "ego"` zooms to one node's neighbourhood, which is
how you inspect a single subject or batch in a graph too large to read whole.

```{r plot-focus, fig.width = 7, fig.height = 4}
plot(graph, focus = "sample_projection", via = c("Subject", "Batch"))
plot(graph, focus = "ego", node = "subject:P1", legend = FALSE)
```

### Reshaping a graph you already have

You do not have to go back to the node and edge sets to change a graph.
`subset_graph()` restricts it to a set of samples and keeps the structure those
samples actually reach; `combine_graphs()` takes the union of several graphs,
rejecting contradictory definitions of the same node; and `add_edges()` appends
an edge set, for example one produced later from a kinship table. Each returns
a new, independently validated graph and leaves its inputs untouched.

```{r graph-edit}
train_graph <- subset_graph(graph, samples = c("S1", "S2", "S3", "S4"))
summary(train_graph)$node_types

# Grouping on the subset graph agrees with asking for a subset directly.
identical(
  grouping_vector(derive_split_constraints(train_graph, mode = "subject")),
  grouping_vector(derive_split_constraints(graph, mode = "subject",
                                           samples = c("S1", "S2", "S3", "S4")))
)
```

`combine_graphs()` is the inverse operation, and the natural way to assemble a
cohort that arrived in pieces — one graph per batch of metadata, merged once
they are all in:

```{r graph-combine}
rejoined <- combine_graphs(
  subset_graph(graph, samples = c("S1", "S2", "S3")),
  subset_graph(graph, samples = c("S4", "S5", "S6")),
  graph_name = "rejoined"
)

identical(summary(rejoined)$node_types, summary(graph)$node_types)
identical(
  grouping_vector(derive_split_constraints(rejoined, mode = "subject")),
  grouping_vector(derive_split_constraints(graph, mode = "subject"))
)
```

`add_edges()` covers the case where a relation only becomes available later.
Kinship, for instance, arrives as its own matrix rather than as a metadata
column; once you have it, thresholding it produces an edge set you can attach to
the graph you already built, and `mode = "relatedness"` becomes available on it:

```{r graph-add-edges}
kinship_table <- data.frame(
  id1     = c("P1", "P2"),
  id2     = c("P2", "P3"),
  kinship = c(0.25, 0.20),
  stringsAsFactors = FALSE
)

graph_with_kinship <- add_edges(
  graph,
  relatedness_edges_from_kinship(kinship_table, threshold = 0.1)
)

summary(graph_with_kinship)$edge_types
grouping_vector(derive_split_constraints(graph_with_kinship, mode = "relatedness"))
```

`P1`, `P2` and `P3` form one relatedness component, so the five samples carried
by those three subjects can no longer be separated; `P4`'s single sample stays
on its own. The thresholded relations are covered in full further down and in the
*modeling-structure* vignette.

### Exporting for other tools

JSON is the interchange contract, but for a look at the graph in Cytoscape,
Gephi, or networkx, `export_graph()` writes GraphML, GML, or flat node and edge
tables. Attributes are flattened into `attr_*` columns, since those formats
cannot hold a nested list.

```{r export, eval = FALSE}
export_graph(graph, "graph.graphml", format = "graphml")
export_graph(graph, "nodes.csv", format = "nodes_csv")
```

## Validate before you split

Validation is where `splitGraph` starts paying off. The graph below is
structurally valid, but it still carries leakage-relevant warnings and
advisories.

```{r validation}
validation <- validate_graph(graph)

validation
as.data.frame(validation)[, c("level", "severity", "code", "message")]
```

That output is the core value proposition of the package in one place:

- repeated subjects are surfaced explicitly
- cross-study subject overlap is surfaced explicitly
- full-dataset feature provenance is surfaced explicitly
- heavy batch reuse is surfaced explicitly

`valid = TRUE` here means the graph has no errors. It does not mean the dataset
is free of leakage risk. Warnings and advisories still matter.

The report is a `depgraph_validation_report`: `print()` gives the severity
roll-up above and `as.data.frame()` the issue table, with one row per finding
and the node and edge ids it implicates.

Findings arrive on three `level`s. **Structural** and **semantic** rules police
the graph itself — dangling edges, duplicate ids, an unsupported relation, a
sample assigned to two subjects, a `time_index` that contradicts the precedence
edges. The **leakage** layer is the one specific to this package, and it is
small enough to list in full:

| `code` | Default severity | Fires when |
|---|---|---|
| `repeated_subject_samples` | advisory | a subject has more than one sample |
| `subject_cross_study_overlap` | warning | one subject's samples span several studies |
| `subject_cross_site_overlap` | warning | one subject's samples span several sites |
| `per_dataset_featureset` | advisory | a `FeatureSet` node declares `derivation_scope = "per_dataset"` |
| `shared_featureset_provenance` | advisory | one feature set is used by more than one sample |
| `missing_time_ordering` | warning | timepoint-linked samples exist with neither `time_index` nor `timepoint_precedes` |
| `heavy_batch_reuse` | advisory | one batch holds at least half the cohort (minimum 3 samples) |

None of these are errors, because none of them is wrong on its own — a
longitudinal study is *supposed* to repeat subjects. They are the inputs to the
split decision, which is why the report describes and does not prescribe.

Pass `levels = "leakage"` to run only that layer, and `severities =` to narrow
the returned table further:

```{r validation-filter}
as.data.frame(
  validate_graph(graph, levels = "leakage", severities = "warning")
)[, c("severity", "code", "message")]
```

The package is also intentionally strict about silent failure. If you ask for a
subset of samples and some of them do not resolve, it errors instead of
dropping them.

```{r strictness}
tryCatch(
  derive_split_constraints(graph, mode = "subject", samples = c("S1", "BAD")),
  error = function(e) e$message
)
```

That behavior is important in practice because quietly omitting samples would
change the truth of the split problem.

### When you really do need to relax a check

Some leakage-relevant rules have legitimate exceptions. For example, in
some pooled designs a single biological "sample" is intentionally paired
across two subjects. The semantic validator flags this as an error by
default, and `derive_split_constraints(mode = "subject")` refuses to
choose a subject for you — both consistent with the "no silent guessing"
stance.

To make the contrast concrete, build a tiny graph in which sample `S1`
is linked to two subjects:

```{r overrides-fixture}
multi_nodes <- graph_node_set(data.frame(
  node_id   = c("sample:S1", "subject:P1", "subject:P2"),
  node_type = c("Sample", "Subject", "Subject"),
  node_key  = c("S1", "P1", "P2"),
  label     = c("S1", "P1", "P2"),
  attrs     = I(list(list(), list(), list())),
  stringsAsFactors = FALSE
))

multi_edges <- graph_edge_set(data.frame(
  edge_id   = c("sample_belongs_to_subject:1", "sample_belongs_to_subject:2"),
  from      = c("sample:S1", "sample:S1"),
  to        = c("subject:P1", "subject:P2"),
  edge_type = c("sample_belongs_to_subject", "sample_belongs_to_subject"),
  attrs     = I(list(list(), list())),
  stringsAsFactors = FALSE
))

multi_graph <- dependency_graph(nodes = multi_nodes, edges = multi_edges)
```

Default behavior — the validator surfaces the multi-subject sample as an
error, and constraint derivation refuses to pick a subject:

```{r overrides-default}
default_report <- validate_graph(multi_graph)
default_report$valid
default_report$issues[, c("severity", "code", "message")]

tryCatch(
  derive_split_constraints(multi_graph, mode = "subject"),
  error = function(e) e$message
)
```

When the ambiguity is intended, opt in explicitly via
`validation_overrides`. The same key
(`allow_multi_subject_samples = TRUE`) is honored by both the validator
and the constraint deriver:

```{r overrides-on}
# Validator: pass.
permissive_report <- validate_graph(
  multi_graph,
  validation_overrides = list(allow_multi_subject_samples = TRUE)
)
permissive_report$valid

# Constraint derivation: pick the first listed subject and record the
# ambiguity in metadata$warnings instead of erroring.
multi_graph$metadata$validation_overrides <-
  list(allow_multi_subject_samples = TRUE)
relaxed_constraint <- derive_split_constraints(multi_graph, mode = "subject")
relaxed_constraint$sample_map[, c("sample_id", "group_id")]
relaxed_constraint$metadata$warnings
```

The override is documented under `?validate_graph`. Use it sparingly and
only when the relaxation matches your scientific intent — the message in
`metadata$warnings` is the audit trail that the choice was tolerated, not
hidden.

## Query the graph to inspect hidden structure

Seven query functions read the graph without changing it. They all return a
`graph_query_result`, so `as.data.frame()` gives you a tidy table in every case.
Five of them work on the typed graph itself:

| Function | Answers |
|---|---|
| `query_node_type()` | which nodes of a given type exist |
| `query_edge_type()` | which edges of a given relation exist, optionally around given nodes |
| `query_neighbors()` | what a node is directly attached to |
| `query_paths()` | *every* simple route between two nodes |
| `query_shortest_paths()` | the shortest such route |

and two project the graph down to samples — `detect_shared_dependencies()` and
`detect_dependency_components()` — which is what the splitting question actually
asks.

Start with the inventory queries. They are the fastest way to check that the
graph contains what you think it contains:

```{r inventory-queries}
as.data.frame(query_node_type(graph, "Batch"))[, c("node_id", "node_type", "node_key")]

# Which samples went through batch B1?
as.data.frame(
  query_edge_type(graph, "sample_processed_in_batch", node_ids = "batch:B1")
)[, c("edge_id", "from", "to")]
```

Then inspect local provenance and trace paths.

```{r neighbors-and-paths}
neighbors_s1 <- query_neighbors(graph, node_ids = "sample:S1", direction = "out")
neighbors_s1
as.data.frame(neighbors_s1)[, c("seed_node_id", "node_id", "node_type", "edge_type")]

subject_outcome_path <- query_shortest_paths(
  graph,
  from = "sample:S1",
  to = "outcome:O_case",
  edge_types = c("sample_belongs_to_subject", "subject_has_outcome")
)

subject_outcome_path
as.data.frame(subject_outcome_path)
```

The first query shows everything the graph knows directly about `S1`. The
second shows that `S1` reaches the subject-level outcome through its subject
node, which is exactly the kind of relationship that would stay implicit in a
plain metadata table.

`query_shortest_paths()` returns one route. When the question is *how many
different ways* two samples are entangled, use `query_paths()`, which enumerates
every simple path. Setting `mode = "all"` ignores edge direction — dependency
edges point from the sample outwards, so any sample-to-sample route has to
traverse one of them backwards — and `max_length = 2` restricts the answer to
dependencies shared one hop away:

```{r all-paths}
s1_to_s6 <- query_paths(graph, from = "sample:S1", to = "sample:S6",
                        mode = "all", max_length = 2)
s1_to_s6
as.data.frame(s1_to_s6)[, c("path_id", "step", "node_id", "node_type")]
```

Three separate routes connect `S1` and `S6`: the batch, the assay, and the
feature set. A single shared column would have shown you one of them.
`max_length` defaults to a finite cap of 8 edges so that a dense graph cannot
make the enumeration explode; when a returned path sits at the cap, the result
records `metadata$truncated = TRUE` to say longer routes may have been
suppressed. Pass `max_length = Inf` to search exhaustively.

```{r projected-dependencies}
shared_dependencies <- detect_shared_dependencies(
  graph,
  via = c("Subject", "Batch", "FeatureSet")
)

as.data.frame(shared_dependencies)[, c(
  "sample_id_1", "sample_id_2", "shared_node_type", "shared_node_id", "edge_type"
)]

dependency_components <- detect_dependency_components(
  graph,
  via = c("Subject", "Batch")
)

as.data.frame(dependency_components)
```

These projected queries are useful because they answer the splitting question
directly. They tell you which samples should be treated as structurally linked,
not just which metadata columns happen to match.

## Derive split constraints from the graph

`splitGraph` derives *direct* constraints — one grouping node per sample — for
`subject`, `batch`, `study`, `time`, `site`, `region`, `platform`, and `assay`,
plus *composite* constraints that combine several dependency sources, and
*pairwise* constraints (`relatedness`, `spatial`) built from thresholded
similarity. This section demonstrates the core four (`subject`, `batch`,
`study`, `time`) and both composite strategies; the cluster relations (`site`,
`region`, `platform`, `assay`) and the pairwise relations are covered in their
own sections below and, in more depth, in the *modeling-structure* vignette.

```{r constraints}
subject_constraint <- derive_split_constraints(graph, mode = "subject")
batch_constraint <- derive_split_constraints(graph, mode = "batch")
study_constraint <- derive_split_constraints(graph, mode = "study")
time_constraint <- derive_split_constraints(graph, mode = "time")

strict_constraint <- derive_split_constraints(
  graph,
  mode = "composite",
  strategy = "strict",
  via = c("Subject", "Batch")
)

rule_based_constraint <- derive_split_constraints(
  graph,
  mode = "composite",
  strategy = "rule_based",
  priority = c("batch", "study", "subject", "time")
)

constraint_overview <- do.call(rbind, lapply(
  list(
    subject = subject_constraint,
    batch = batch_constraint,
    study = study_constraint,
    time = time_constraint,
    composite_strict = strict_constraint,
    composite_rule = rule_based_constraint
  ),
  function(x) {
    data.frame(
      strategy = x$strategy,
      groups = length(unique(x$sample_map$group_id)),
      warnings = if (is.null(x$metadata$warnings)) 0L else length(x$metadata$warnings),
      stringsAsFactors = FALSE
    )
  }
))

constraint_overview <- cbind(constraint = row.names(constraint_overview), constraint_overview)
row.names(constraint_overview) <- NULL

constraint_overview
```

That summary already shows why the package is useful: different notions of
dependency produce different splitting units.

### Batch constraints

```{r batch-constraint}
batch_constraint
as.data.frame(batch_constraint)[, c("sample_id", "group_id", "group_label", "explanation")]
```

Batch grouping keeps all `B1` samples together and preserves `S5` as an
explicit singleton because it has no batch assignment. Missing structure is not
hidden.

### Time constraints

```{r time-constraint}
time_constraint
as.data.frame(time_constraint)[, c("sample_id", "group_id", "timepoint_id", "order_rank")]
```

Time grouping adds `order_rank`, which is the field downstream tooling actually
needs for ordered evaluation. The missing timepoint on `S5` stays visible as
`NA`, so ordering is partial rather than pretended.

### Composite constraints

```{r composite-constraints}
strict_constraint
as.data.frame(strict_constraint)[, c("sample_id", "group_id", "constraint_type")]

rule_based_constraint
as.data.frame(rule_based_constraint)[, c("sample_id", "group_id", "constraint_type", "group_label")]
```

The strict composite constraint uses transitive closure: `S1`, `S2`, `S3`, and
`S6` end up in the same group because subject and batch links connect them into
one dependency component. The rule-based composite constraint is different: it
uses the highest-priority available dependency per sample, so `S5` falls back
to study-level grouping instead of becoming a composite component.

## Cluster-style relations: site, region, platform, assay

Beyond subject/batch/study/time, several other metadata columns define
*cluster-style* leakage axes — categorical groupings a model can memorize
instead of generalizing across. `splitGraph` models each as a first-class node
type with its own auto-detected column, validation rule, and constraint mode:

| Relation | Column | Edge | Mode |
|---|---|---|---|
| Collection site / center | `site_id` | `sample_collected_at_site` | `"site"` |
| Tissue / anatomical region | `region_id` | `sample_located_in_region` | `"region"` |
| Sequencing / measurement platform | `platform_id` | `sample_run_on_platform` | `"platform"` |
| Assay / modality | `assay_id` | `sample_measured_by_assay` | `"assay"` |

They all behave identically: `graph_from_metadata()` auto-detects the column,
`validate_graph()` flags samples assigned to more than one target, the mode
groups samples so no cluster straddles a split, and `as_split_spec()` carries
the assignment as a blocking annotation (`site_group`, `region_group`,
`platform_group`, `assay_group`).

### Site

In multi-site studies, the collection site is a common leakage axis: a model can
"recognize" a site rather than generalize across sites. `splitGraph` models the
site as a first-class `Site` node connected by `sample_collected_at_site` edges.
`graph_from_metadata()` auto-detects a `site_id` column, so no extra wiring is
needed.

```{r site-structure}
site_meta <- data.frame(
  sample_id  = c("S1", "S2", "S3", "S4", "S5", "S6"),
  subject_id = c("P1", "P1", "P2", "P2", "P3", "P3"),
  site_id    = c("NYC", "NYC", "BOS", "BOS", "NYC", "BOS"),
  stringsAsFactors = FALSE
)

site_graph <- graph_from_metadata(site_meta, graph_name = "multi-site")

# Group samples so that no collection site straddles a train/test split.
site_constraint <- derive_split_constraints(site_graph, mode = "site")
as.data.frame(site_constraint)[, c("sample_id", "group_id", "group_label")]
```

`mode = "site"` keeps every sample from a given site in the same group. The site
assignment is also carried into the `split_spec` as a blocking annotation
(`site_group`), so a downstream consumer can block on site even when the primary
grouping is something else (e.g. subject):

```{r site-spec}
subject_then_block_by_site <- as_split_spec(
  derive_split_constraints(site_graph, mode = "subject"),
  graph = site_graph
)
subject_then_block_by_site$block_vars
head(subject_then_block_by_site$sample_data[, c("sample_id", "group_id", "site_group")])
```

Samples assigned to more than one site are rejected rather than silently
resolved, by both `validate_graph()` and `derive_split_constraints(mode = "site")`.

Validation also watches the other direction. Subject `P3` above contributed `S5`
at `NYC` and `S6` at `BOS`, so a site-grouped split would separate two samples
from the same person — the mirror image of the cross-study overlap seen earlier,
and the reason to read the validation report before committing to a mode:

```{r site-validation}
as.data.frame(validate_graph(site_graph))[
  , c("severity", "code", "message")
]
```

### Region

The `Region` relation works the same way for a categorical tissue or anatomical
region (`region_id` column, `sample_located_in_region` edge,
`derive_split_constraints(mode = "region")`), so samples from the same region
stay together across a split.

```{r region-structure}
region_meta <- data.frame(
  sample_id = c("S1", "S2", "S3", "S4"),
  region_id = c("cortex", "cortex", "hippocampus", "hippocampus"),
  stringsAsFactors = FALSE
)
region_graph <- graph_from_metadata(region_meta, graph_name = "regions")
as.data.frame(derive_split_constraints(region_graph, mode = "region"))[
  , c("sample_id", "group_id", "group_label")
]
```

### Platform and assay

Technical measurement structure is the same story. `platform_id` (the
sequencing instrument or measurement platform) and `assay_id` (the assay or
modality) are both auto-detected, and `mode = "platform"` / `mode = "assay"`
group samples so a whole platform or assay never straddles a split — useful when
a batch effect tracks the instrument rather than the run.

```{r platform-assay}
tech_meta <- data.frame(
  sample_id   = c("S1", "S2", "S3", "S4"),
  platform_id = c("illumina", "illumina", "nanopore", "nanopore"),
  assay_id    = c("rnaseq", "rnaseq", "wgs", "wgs"),
  stringsAsFactors = FALSE
)
tech_graph <- graph_from_metadata(tech_meta, graph_name = "tech")

as.data.frame(derive_split_constraints(tech_graph, mode = "platform"))[
  , c("sample_id", "group_id", "group_label")
]
as.data.frame(derive_split_constraints(tech_graph, mode = "assay"))[
  , c("sample_id", "group_id", "group_label")
]
```

These new types are first-class in the typed layout too. Building a graph that
carries several of them shows how each gets its own colour and layer band — the
same visual vocabulary as the core types, so a mixed graph stays readable:

```{r cluster-plot, fig.width = 7, fig.height = 5}
mixed_meta <- data.frame(
  sample_id   = c("S1", "S2", "S3", "S4"),
  subject_id  = c("P1", "P1", "P2", "P2"),
  site_id     = c("NYC", "NYC", "BOS", "BOS"),
  platform_id = c("illumina", "illumina", "nanopore", "nanopore"),
  stringsAsFactors = FALSE
)
plot(graph_from_metadata(mixed_meta, graph_name = "mixed_structure"))
```

### Pairwise relations: relatedness and spatial proximity

Not every leakage source is a clean categorical group. Genetic relatedness and
spatial proximity are *pairwise and continuous*: they link individual pairs by a
similarity score. `splitGraph` models these as thresholded, undirected edges
built with `relatedness_edges_from_kinship()` and `spatial_edges_from_coords()`,
then `mode = "relatedness"` / `mode = "spatial"` form groups by transitive
closure over the surviving edges — so a chain of individually near neighbours
still lands in one group, a grouping a single column cannot express.

The example below keeps subject pairs whose kinship is at least `0.1`. `P1`–`P2`
and `P2`–`P3` clear the threshold, so those three subjects collapse into one
group by transitive closure; the unrelated `P4` stays on its own:

```{r relatedness-demo}
kin <- data.frame(
  id1     = c("P1", "P2", "P1"),
  id2     = c("P2", "P3", "P4"),
  kinship = c(0.25, 0.20, 0.02),   # P1-P4 is below the 0.1 threshold
  stringsAsFactors = FALSE
)
rel_edges <- relatedness_edges_from_kinship(kin, threshold = 0.1)

rel_meta <- data.frame(
  sample_id  = paste0("S", 1:4),
  subject_id = c("P1", "P2", "P3", "P4"),
  stringsAsFactors = FALSE
)
rel_graph <- build_dependency_graph(
  nodes = list(
    create_nodes(rel_meta, "Sample", "sample_id"),
    create_nodes(rel_meta, "Subject", "subject_id")
  ),
  edges = list(
    create_edges(rel_meta, "sample_id", "subject_id",
                 "Sample", "Subject", "sample_belongs_to_subject"),
    rel_edges
  )
)

grouping_vector(derive_split_constraints(rel_graph, mode = "relatedness"))
```

The `spatial` mode works identically over sample coordinates
(`spatial_edges_from_coords(coords, radius)`). Both relations — with the full
scikit-learn handoff — are covered in depth in the *modeling-structure*
vignette:

```{r modeling-pointer, eval = FALSE}
vignette("modeling-structure", package = "splitGraph")
```

## Time ordering can come from precedence edges alone

If explicit `time_index` metadata are unavailable, `splitGraph` can still infer
time order from `timepoint_precedes` edges.

```{r precedence-only}
precedence_meta <- data.frame(
  sample_id = c("S1", "S2", "S3"),
  subject_id = c("P1", "P1", "P2"),
  study_id = c("ST1", "ST1", "ST2"),
  timepoint_id = c("T0", "T1", "T2"),
  stringsAsFactors = FALSE
)

precedence_graph <- build_dependency_graph(
  nodes = list(
    create_nodes(precedence_meta, type = "Sample", id_col = "sample_id"),
    create_nodes(precedence_meta, type = "Subject", id_col = "subject_id"),
    create_nodes(precedence_meta, type = "Study", id_col = "study_id"),
    create_nodes(
      data.frame(timepoint_id = c("T0", "T1", "T2"), stringsAsFactors = FALSE),
      type = "Timepoint",
      id_col = "timepoint_id"
    )
  ),
  edges = list(
    create_edges(
      precedence_meta, "sample_id", "subject_id",
      "Sample", "Subject", "sample_belongs_to_subject"
    ),
    create_edges(
      precedence_meta, "sample_id", "study_id",
      "Sample", "Study", "sample_from_study"
    ),
    create_edges(
      precedence_meta, "sample_id", "timepoint_id",
      "Sample", "Timepoint", "sample_collected_at_timepoint"
    ),
    create_edges(
      data.frame(
        from_timepoint = c("T0", "T1"),
        to_timepoint = c("T1", "T2"),
        stringsAsFactors = FALSE
      ),
      "from_timepoint", "to_timepoint",
      "Timepoint", "Timepoint", "timepoint_precedes"
    )
  ),
  graph_name = "precedence_only_graph"
)

precedence_time_constraint <- derive_split_constraints(precedence_graph, mode = "time")

precedence_time_constraint$metadata$time_order_source
as.data.frame(precedence_time_constraint)[, c("sample_id", "timepoint_id", "time_index", "order_rank")]
```

The important detail is that ordering is still derived, but the source is
`timepoint_precedes` rather than `time_index`.

## Translate the constraint into a split specification

The graph-derived constraint is not the end of the workflow. The main handoff
target is a canonical sample-level split specification — the `split_spec`
class. Downstream tools consume it through their own adapters, so
`split_spec` stays tool-agnostic.

```{r split-spec}
split_spec <- as_split_spec(strict_constraint, graph = graph)
split_spec

as.data.frame(split_spec)[, c(
  "sample_id", "group_id", "batch_group", "study_group",
  "stratum", "timepoint_id", "order_rank"
)]

split_spec_validation <- validate_split_spec(split_spec)
split_spec_validation
as.data.frame(split_spec_validation)
```

This translation step is where the package becomes operational for downstream
evaluation workflows:

- `group_id` carries the split unit
- `batch_group` and `study_group` are available for blocking, alongside
  `site_group`, `region_group`, `platform_group` and `assay_group` when the
  graph carries those relations
- `order_rank` is available for ordered evaluation
- `stratum` carries the outcome level each sample has, so a consumer can
  stratify; it is an annotation only, and `splitGraph` never balances folds
- the generated object is validated before handoff

The declared roles are readable off the spec itself, which is what an adapter
keys on rather than guessing column names:

```{r spec-roles}
split_spec$group_var
split_spec$block_vars
split_spec$time_var
split_spec$stratum_var
```

## Summarize the leakage picture in one object

The final helper combines graph validation, constraint diagnostics, and
split-spec readiness into one `leakage_risk_summary` object. Crucially, when
you pass the `constraint` you chose, it reports a `severed` column: whether that
constraint **structurally eliminates** each leakage path (`TRUE`), leaves it
open (`FALSE`), or is not applicable (`NA`, e.g. for informational split-spec
rows).

```{r risk-summary}
risk_summary <- summarize_leakage_risks(
  graph,
  constraint = strict_constraint,
  split_spec = split_spec
)

risk_summary
as.data.frame(risk_summary)[, c("source", "severity", "category", "severed", "message")]
```

Read the `severed` column against the constraint you actually chose. Here the
strict composite constraint (`via = c("Subject", "Batch")`) severs the subject,
cross-study, and batch-reuse risks (`TRUE`) because those samples are forced
into the same group — but it does **not** address `per_dataset_featureset` or
`shared_featureset_provenance` (`FALSE`), which are orthogonal to a
subject/batch grouping: every sample shares `FS_GLOBAL`, so no grouping of
samples can undo a feature set that was fitted on the whole cohort. The
`NA` rows are the constraint and split-spec diagnostics, which describe
readiness rather than a leakage path. That distinction is the point: the
summary tells you which of your surfaced risks your split design has actually
handled, and which still need attention (a different mode, a blocking variable,
or a data fix) — so the leakage trade-off is explicit before any model is
trained.

## Downstream handoff

`split_spec` is the tool-agnostic handoff artifact. `splitGraph` does not
know about any particular resampling package — downstream consumers provide
their own adapters so `splitGraph` stays neutral. The typical end-to-end
flow is:

1. `graph_from_metadata()` (or the explicit constructor path) → typed
   `dependency_graph`
2. `derive_split_constraints(g, mode = ...)` → `split_constraint`
3. `as_split_spec(constraint, graph = g)` → `split_spec`
4. (optional) `write_split_spec(spec, path)` → JSON, for cross-session or
   cross-language handoff
5. adapter in the downstream package → native resamples

The `sample_data` frame carried by `split_spec` exposes exactly the columns
downstream adapters consume: `sample_id` for joining against the observation
frame, `group_id` for grouped resampling, `batch_group` / `study_group` for
blocking, and `order_rank` for ordered evaluation. Adapters can be built by
any package that wants to consume a `split_spec` — for example, on top of
`rsample::group_vfold_cv()` (grouped CV keyed to `group_id`) or
`rsample::rolling_origin()` (ordered evaluation keyed to `order_rank`).

### Persisting the handoff

If the downstream consumer is in a different R session — or in a different
language entirely — write the spec (and, if useful, the graph) to JSON. The
on-disk format has a formal JSON Schema (Draft 2020-12) shipped under
`inst/schema/<schema_version>/`, and each written file references its own
version via a `$schema` key, so the reference stays valid after later schema
bumps. You can validate a handoff file against that contract with
`validate_split_spec_json()` before consuming it, or ask the reader to do it by
passing `validate = TRUE`, and `NA` values round-trip as JSON `null`.

```{r serialize, eval = requireNamespace("jsonlite", quietly = TRUE)}
spec_path <- tempfile(fileext = ".json")
write_split_spec(split_spec, spec_path)

# Validate the file against the shipped JSON Schema.
validate_split_spec_json(spec_path)$valid

# Round-trip it back into R unchanged. `validate = TRUE` checks the file
# against the shipped schema before parsing and runs the preflight validator
# on the result, failing with a classed error instead of a silent surprise.
spec_round_trip <- read_split_spec(spec_path, validate = TRUE)
identical(split_spec$sample_data$group_id, spec_round_trip$sample_data$group_id)

unlink(spec_path)
```

The graph itself serialises the same way, which is what you want when the
consumer needs the *reasoning* and not only the answer — an audit trail, a
reviewer re-deriving a different mode, or a pipeline that builds the graph once
and derives several constraints later. `write_dependency_graph()` /
`read_dependency_graph()` are the pair, and `validate_graph_json()` is the
file-level check:

```{r serialize-graph, eval = requireNamespace("jsonlite", quietly = TRUE)}
graph_path <- tempfile(fileext = ".json")
write_dependency_graph(graph, graph_path)

validate_graph_json(graph_path)$valid

graph_round_trip <- read_dependency_graph(graph_path, validate = TRUE)

# The typed structure survives exactly ...
identical(as.data.frame(graph$edges), as.data.frame(graph_round_trip$edges))
identical(summary(graph), summary(graph_round_trip))

# ... and so does everything derived from it.
identical(
  grouping_vector(derive_split_constraints(graph_round_trip, mode = "subject")),
  grouping_vector(subject_constraint)
)
identical(
  as.data.frame(validate_graph(graph)),
  as.data.frame(validate_graph(graph_round_trip))
)
```

One JSON detail is worth knowing: the free-form `attrs` bag on each node is
written with ordinary JSON semantics, so an attribute that was `NA` in R comes
back as `NULL` rather than `NA`. Both mean "not recorded" and nothing in
`splitGraph` reads `attrs` to derive a constraint — node and edge types,
identifiers and relations, which is what the split logic uses, round-trip
unchanged.

Both writers stamp the file with the schema version they were produced under and
a `$schema` URL that points at that exact version, so the reference keeps
resolving after a later schema bump:

```{r schema-stamp, eval = requireNamespace("jsonlite", quietly = TRUE)}
on_disk <- jsonlite::fromJSON(graph_path, simplifyVector = FALSE)
on_disk$schema_version
sub(".*/schema/", "", on_disk[["$schema"]])
```

That stamp is what makes an old file readable by a newer splitGraph. When a file
predates the installed schema, `migrate_dependency_graph_json()` and
`migrate_split_spec_json()` rewrite it at the current version, filling anything
introduced since with its default — a missing `split_spec` column becomes `NA`
rather than an error. Files already current are rewritten unchanged, so the call
is safe to run unconditionally on an archive:

```{r migrate, eval = requireNamespace("jsonlite", quietly = TRUE)}
migrated <- migrate_dependency_graph_json(graph_path, tempfile(fileext = ".json"))
validate_graph_json(migrated)$valid

unlink(c(graph_path, migrated))
```

Because the format is a documented, versioned contract, consumers are not
limited to R. The package ships a pure-Python reference reader (`inst/python`)
that recovers the same grouping and ordering and drives scikit-learn resamplers;
the *cross-language-handoff* vignette walks the full R → JSON → Python →
scikit-learn path. For worked R adapters (a base-R leave-one-group-out adapter
and illustrative `rsample` adapters), see the *adapter-cookbook* vignette:

```{r cookbook-pointer, eval = FALSE}
vignette("adapter-cookbook", package = "splitGraph")       # R adapters
vignette("cross-language-handoff", package = "splitGraph") # Python / sklearn
```

## Case studies

The end-to-end workflow above shows the package surface. The case studies below
show how the same graph leads to different evaluation decisions depending on
the scientific question.

### Case study 1: repeated subjects in a longitudinal cohort

Suppose the real question is whether future observations from the same subject
should be held out from training. In this setting, subject reuse and time
ordering both matter, but they solve different problems.

```{r case-study-1}
subject_groups <- grouping_vector(subject_constraint)
time_groups <- time_constraint$sample_map[, c("sample_id", "group_id", "timepoint_id", "order_rank")]

subject_groups
time_groups
```

Interpretation:

- `S1` and `S2` share subject `P1`, so subject-grouped evaluation keeps them
  together.
- `S3` and `S6` share subject `P2`, so they also stay together under a
  subject-based split.
- time grouping adds a different axis: `T0`, `T1`, and `T2` become ordered
  units with explicit `order_rank`.

If the leakage concern is repeated measurements from the same individual, use
the subject constraint. If the evaluation question is prospective prediction,
the time constraint adds the ordering information you need.

### Case study 2: a subject reused across studies

The graph intentionally includes subject `P2` in both `ST1` and `ST2`. A
study-only split would treat those studies as separate units, but the graph
shows that subject overlap breaks the intended independence.

```{r case-study-2}
cross_study_issues <- as.data.frame(validation)[
  as.data.frame(validation)$code == "subject_cross_study_overlap",
  c("severity", "code", "message")
]

p2_shared <- detect_shared_dependencies(
  graph,
  via = "Subject",
  samples = c("S3", "S6")
)

study_only_map <- study_constraint$sample_map[, c("sample_id", "group_id", "group_label")]
strict_map <- strict_constraint$sample_map[, c("sample_id", "group_id", "constraint_type")]

cross_study_issues
as.data.frame(p2_shared)
study_only_map[study_only_map$sample_id %in% c("S3", "S6"), ]
strict_map[strict_map$sample_id %in% c("S3", "S6"), ]
```

Interpretation:

- validation surfaces the cross-study subject overlap directly
- the shared-dependency query confirms that `S3` and `S6` are linked through
  the same subject
- a study-only split would place them in different groups (`ST1` versus `ST2`)
- the strict composite constraint correctly keeps them in the same dependency
  component

This is exactly the kind of failure mode `splitGraph` is designed to expose:
metadata columns suggest a legitimate study split, but graph structure shows
that the split would still leak subject information.

### Case study 3: partially observed technical metadata

Real metadata are rarely complete. Here, `S5` has no batch assignment and no
timepoint assignment. The package does not pretend those fields exist. It keeps
the sample visible and tells you how the split logic handled it.

```{r case-study-3}
batch_missing <- batch_constraint$sample_map[
  batch_constraint$sample_map$sample_id == "S5",
  c("sample_id", "group_id", "group_label", "explanation")
]

rule_based_missing <- rule_based_constraint$sample_map[
  rule_based_constraint$sample_map$sample_id == "S5",
  c("sample_id", "group_id", "constraint_type", "group_label", "explanation")
]

split_spec_missing <- as.data.frame(split_spec)[
  as.data.frame(split_spec)$sample_id == "S5",
  c("sample_id", "group_id", "batch_group", "study_group", "timepoint_id", "order_rank")
]

batch_missing
rule_based_missing
split_spec_missing
```

Interpretation:

- batch-based splitting keeps `S5` as an explicit singleton because batch
  metadata are missing
- the rule-based composite strategy falls back to study-level grouping for `S5`
- the translated split specification preserves the missing batch and time
  fields as `NA` rather than silently inventing values

That behavior matters because incomplete metadata are common. `splitGraph`
stays strict about what is known, but still produces a usable, inspectable
split object.

### Case study 4: choosing a defensible split strategy

A typical practical question is not "what can the package compute?" but "which
constraint should I actually use?" The answer depends on which dependency
source is scientifically unacceptable to leak across train and test.

```{r case-study-4}
strategy_summary <- data.frame(
  constraint = c("subject", "batch", "study", "time", "composite_strict", "composite_rule"),
  groups = c(
    length(unique(subject_constraint$sample_map$group_id)),
    length(unique(batch_constraint$sample_map$group_id)),
    length(unique(study_constraint$sample_map$group_id)),
    length(unique(time_constraint$sample_map$group_id)),
    length(unique(strict_constraint$sample_map$group_id)),
    length(unique(rule_based_constraint$sample_map$group_id))
  ),
  warnings = c(
    length(or_empty(subject_constraint$metadata$warnings)),
    length(or_empty(batch_constraint$metadata$warnings)),
    length(or_empty(study_constraint$metadata$warnings)),
    length(or_empty(time_constraint$metadata$warnings)),
    length(or_empty(strict_constraint$metadata$warnings)),
    length(or_empty(rule_based_constraint$metadata$warnings))
  ),
  recommended_resampling = c(
    as_split_spec(subject_constraint, graph = graph)$recommended_resampling,
    as_split_spec(batch_constraint, graph = graph)$recommended_resampling,
    as_split_spec(study_constraint, graph = graph)$recommended_resampling,
    as_split_spec(time_constraint, graph = graph)$recommended_resampling,
    as_split_spec(strict_constraint, graph = graph)$recommended_resampling,
    as_split_spec(rule_based_constraint, graph = graph)$recommended_resampling
  ),
  stringsAsFactors = FALSE
)

strategy_summary
```

Interpretation:

- subject grouping is the right default when repeated individuals are the
  dominant leakage source
- batch grouping is appropriate when technical runs are the main contamination
  risk
- study grouping is useful for cross-study generalization only when no higher
  level dependency crosses study boundaries
- strict composite grouping is the safest choice when multiple dependency
  sources can connect samples transitively
- rule-based composite grouping is a pragmatic fallback when you want a single
  deterministic hierarchy over partially observed metadata

The package does not choose the scientific objective for you. It makes the
trade-off visible and auditable.

## When `splitGraph` is useful

`splitGraph` is a good fit when:

- sample relationships are scientifically meaningful and must influence
  evaluation
- metadata contain repeated subjects, shared batches, multiple studies, or
  temporal structure
- feature provenance or outcome level matters for leakage assessment
- you want deterministic, inspectable split constraints instead of ad hoc
  grouping code

## What `splitGraph` is not for

`splitGraph` is not:

- a general biological network analysis package
- a model training framework
- a resampling engine
- a substitute for downstream performance auditing

Its value is earlier in the workflow: it makes dependency structure explicit so
that the split design itself can be justified.

## Takeaway

If you already know your data have repeated subjects, reused batches, temporal
ordering, or shared feature provenance, then you already have a graph problem
whether you model it explicitly or not. `splitGraph` is useful because it turns
that hidden graph into an object you can validate, query, and convert into a
split design that downstream tooling can trust.
