# Scientific Method Selection

Read this catalog once when a new CATAIL goal begins or when its question, available evidence, or requested output materially changes. It helps select the smallest sufficient combination of research modes. It is not a checklist requiring every mode.

## Selection rule

Choose methods by asking:

1. What knowledge change is sought: description, explanation, prediction, causal effect, interpretation, validation, synthesis, or artifact performance?
2. What is currently available: only an idea, prior literature, naturally occurring data, manipulable conditions, computational systems, human accounts, an artifact, or a published result?
3. Which serious rival explanations must the evidence distinguish?
4. What can the selected method establish, and what remains outside its inferential reach?
5. What is the cheapest ethical design that can materially update the claim?
6. What result would support, contradict, leave unresolved, or invalidate the attempted test?

Record the chosen modes, why each is necessary, what it contributes, and why omitted modes are unnecessary. Prefer information gain over procedural completeness.

## 1. Evidence Synthesis

**Use when**

- mapping a field, testing candidate novelty, resolving disagreement, identifying boundary conditions, or answering a question primarily from existing studies;
- preparing a narrative, scoping, systematic, rapid, or meta-analytic review;
- a new empirical study must be positioned against inspected prior work.

**Starting conditions**

- an explicit question or concept map;
- defined source types, databases, languages, dates, screening scope, and stopping rule;
- a declared synthesis type proportional to the user's goal.

**Produces**

- a source-level map of claims, methods, populations/systems, results, quality, disagreement, and missing evidence;
- bounded estimates or qualitative synthesis when studies are sufficiently compatible;
- candidate novelty and contribution statements, never proof of absence.

**Cannot establish alone**

- that an unobserved study does not exist;
- a causal or mechanistic conclusion stronger than the included evidence;
- empirical validity for a new system merely because prior work is favorable.

**Validity risks and quality floor**

- terminology blind spots, database coverage, publication bias, duplicate reports, selective screening, incompatible estimands, and abstract-only interpretation;
- retain exact queries and source identifiers; distinguish located, abstract-screened, and full-text-inspected sources; assess quality and contradictory evidence; keep synthesis boundaries visible.

**Common combinations**

- Theory Building for mechanism and rival models;
- Experimental, Observational, or Computational Research for a new unresolved question;
- Reproduction & Replication when influential evidence is fragile.

**Exit**

- stop when the declared coverage and saturation rule is met, then report what was and was not searched. Route to positioning, synthesis output, or another method only if the evidence map justifies it.

## 2. Exploratory Discovery

**Use when**

- the phenomenon, variables, categories, mechanisms, or useful hypotheses are not yet well formed;
- unexpected patterns or anomalies deserve structured investigation;
- a pilot is needed to validate measurement, task construction, or feasibility.

**Starting conditions**

- a bounded exploration space and data provenance;
- a clear distinction between observations available before exploration and patterns discovered after inspection;
- a stopping or saturation rule that is not “continue until interesting.”

**Produces**

- candidate patterns, constructs, questions, mechanisms, subgroups, measurements, or hypotheses;
- feasibility and measurement knowledge;
- exploratory evidence that can motivate independent confirmation.

**Cannot establish alone**

- prespecified confirmation on the same target data;
- reliable effect magnitude after unrestricted search;
- generality beyond the explored sample or system.

**Validity risks and quality floor**

- data dredging, multiple comparisons, researcher degrees of freedom, leakage, unstable clusters, hindsight bias, and narrative overfitting;
- label outputs exploratory; preserve all materially attempted analyses; use resampling or held-out data when appropriate; register result-derived hypotheses as new candidates.

**Common combinations**

- Theory Building to turn patterns into mechanisms and discriminating predictions;
- Experimental or Observational Research on new data;
- Qualitative Methods to interpret unexpected categories or experiences.

**Exit**

- stop when exploration yields a bounded candidate, a diagnosed measurement problem, saturation, or low information gain. Confirmation requires a new design and independent target evidence.

## 3. Theory Building

**Use when**

- proposing or refining a mechanism, conceptual model, taxonomy, formal account, or explanatory framework;
- several observations or literatures need a coherent explanation;
- rival explanations must be made distinguishable before investigation.

**Starting conditions**

- clearly defined constructs and the observations or problems the theory addresses;
- nearest theories or explanations and their known limits;
- a scope in which the theory is intended to operate.

**Produces**

- explicit constructs, relations, assumptions, boundary conditions, causal or logical structure, and discriminating predictions;
- rival models and observations that would weaken or revise them;
- conceptual integration or formal propositions.

**Cannot establish alone**

- that the favored mechanism operates in the world;
- empirical superiority without discriminating evidence;
- novelty without prior-art positioning.

**Validity risks and quality floor**

- unfalsifiability, circular definitions, relabeling rather than explanation, post-hoc accommodation, hidden assumptions, and predictions shared by all rivals;
- define constructs operationally where possible; state scope and assumptions; derive risky predictions; identify observations compatible with the theory but not diagnostic of it.

**Common combinations**

- Evidence Synthesis for prior theories and empirical constraints;
- Experimental or Observational Research for discriminating tests;
- Computational Research for simulation, formal consequences, or model comparison.

**Exit**

- proceed when the model generates distinct, testable predictions or a defensible conceptual contribution. Reframe when it cannot differ observably from rivals.

## 4. Observational Research

**Use when**

- studying naturally occurring variation, associations, trajectories, prevalence, prediction, or contexts where intervention is infeasible or unethical;
- working with cohorts, logs, archives, surveys, sensors, traces, or administrative data.

**Starting conditions**

- a defined population/system, sampling or inclusion mechanism, variables, time structure, and unit of inference;
- a defensible measurement model and data-generation account;
- confounders, selection mechanisms, missingness, and plausible causal structures identified when causal language is contemplated.

**Produces**

- descriptive distributions, associations, predictive relationships, temporal patterns, and bounded causal estimates only under explicit identifying assumptions;
- evidence about external conditions difficult to reproduce experimentally.

**Cannot establish alone**

- causality merely from correlation or temporal order;
- mechanism from predictive performance;
- population generality when selection is unknown.

**Validity risks and quality floor**

- confounding, reverse causality, collider bias, selection, measurement error, dependence, survivorship, missingness, and leakage;
- separate descriptive, predictive, and causal estimands; justify adjustment; test sensitivity to plausible violations; align conclusions with the actual sampling and inference units.

**Common combinations**

- Theory Building for causal structure and rival explanations;
- Experimental Research for mechanisms or interventions;
- Qualitative Methods for construct meaning and contextual interpretation.

**Exit**

- conclude within supported estimands and assumptions. Route unresolved causality to a stronger design rather than upgrading language.

## 5. Experimental Research

**Use when**

- estimating intervention effects, testing discriminating predictions, validating mechanisms through controlled manipulation, or comparing conditions under controlled assignment;
- randomized, quasi-experimental, laboratory, field, online, or single-case designs are appropriate.

**Starting conditions**

- a positioned question, explicit hypotheses and rivals, manipulable conditions, valid outcomes, true units of assignment and inference, and an ethical feasible intervention;
- a design capable of detecting a meaningful effect or supporting an equivalence/non-inferiority claim when that is the target.

**Produces**

- estimates of controlled contrasts with uncertainty;
- causal evidence within the design's identification and implementation boundaries;
- mechanism evidence only when mediators, interventions, and alternative pathways are appropriately tested.

**Cannot establish alone**

- unrestricted external validity;
- theory uniqueness when rivals predict the same contrast;
- absence of an effect from a merely non-significant underpowered result.

**Validity risks and quality floor**

- weak controls, pseudoreplication, noncompliance, attrition, demand effects, instrumentation drift, multiplicity, outcome switching, and implementation differences;
- freeze confirmatory choices before target results; use relevant baselines, negative controls, manipulation checks, effect sizes, uncertainty, and transparent exclusions; preserve failed and adverse runs.

**Common combinations**

- Evidence Synthesis and Theory Building before design;
- Computational Research for agent, simulation, or benchmark experiments;
- Reproduction & Replication before broad generalization.

**Exit**

- stop at the planned rule, classify support, contradiction, mixed evidence, inconclusiveness, or method failure, and route any new result-derived hypothesis to a new study.

## 6. Computational Research

**Use when**

- investigating algorithms, models, simulations, agents, software systems, benchmarks, digital experiments, formal computation, or large-scale data processing;
- computation is the research object, measurement instrument, or test environment.

**Starting conditions**

- defined system versions, code and data provenance, environment, task construction, baselines, metrics, seeds, compute budget, and inference unit;
- an explicit account of what real-world or theoretical claim the computation can support.

**Produces**

- reproducible performance estimates, simulations, ablations, sensitivity results, benchmark comparisons, failure modes, or formal computational consequences;
- mechanistic clues when interventions and measurements isolate a plausible component.

**Cannot establish alone**

- real-world generality from one benchmark or model family;
- mechanism from a component label or ablation alone;
- fair superiority when budgets, prompts, data access, or evaluation differ.

**Validity risks and quality floor**

- data or benchmark leakage, evaluator bias, hidden retries, cherry-picked seeds, unfair compute, version drift, correlated tasks, contamination, and non-independent repetitions;
- freeze configurations for confirmation; record code/data/environment and dirty state; expose all material arms and retries; use task-level or system-level uncertainty rather than treating repeated calls as independent by default.

**Common combinations**

- Experimental Research for controlled comparisons;
- Method & Artifact Research for new tools or benchmarks;
- Reproduction & Replication across implementations, models, datasets, and environments.

**Exit**

- conclude only within documented systems, tasks, budgets, evaluators, and versions. Generalization requires planned boundary tests or replication.

## 7. Qualitative & Mixed Methods

**Use when**

- studying meaning, experience, process, context, implementation, stakeholder reasoning, or constructs not adequately represented by existing measures;
- integrating qualitative explanation with quantitative magnitude or comparison.

**Starting conditions**

- a clear sampling rationale, participant or material boundary, data-generation method, consent and confidentiality plan where applicable, and analytic orientation;
- for mixed methods, an explicit reason integration is needed and where it occurs.

**Produces**

- themes, mechanisms, process accounts, taxonomies, cases, counterexamples, contextual explanations, and integrated interpretations;
- construct refinement and explanations for heterogeneous quantitative results.

**Cannot establish alone**

- statistical prevalence from purposive samples;
- universal saturation or objectivity by declaration;
- causal magnitude without an appropriate design.

**Validity risks and quality floor**

- leading prompts, selective quotation, reflexivity blind spots, premature closure, decontextualization, translation loss, and tokenistic mixed-method integration;
- document sampling and analytic decisions; preserve negative cases; distinguish participant accounts from researcher interpretation; use multiple coders or member/context checks only when methodologically justified.

**Common combinations**

- Exploratory Discovery and Theory Building;
- Observational or Experimental Research for triangulation;
- Method & Artifact Research for measurement or intervention design.

**Exit**

- stop at the declared saturation or information-power rule, present counterexamples and reflexive limits, and explain how qualitative and quantitative components change one another.

## 8. Method & Artifact Research

**Use when**

- creating or evaluating a method, instrument, dataset, benchmark, software tool, taxonomy, protocol, or research infrastructure;
- the artifact itself is the scientific contribution.

**Starting conditions**

- a documented unmet research need and nearest alternatives;
- intended users, tasks, construct or target, success dimensions, and adoption or misuse risks;
- a validation plan that goes beyond demonstrating that the artifact runs.

**Produces**

- an artifact plus evidence of validity, reliability, utility, efficiency, robustness, or comparative advantage;
- known failure cases, scope, maintenance, and reproducibility information.

**Cannot establish alone**

- scientific value from implementation novelty;
- construct validity from face validity or benchmark performance alone;
- superiority without relevant baselines and fair conditions.

**Validity risks and quality floor**

- benchmark gaming, circular evaluation, unrepresentative tasks, train/test contamination, missing user evaluation, unstable annotation, and dependency on hidden infrastructure;
- validate intended constructs and uses; compare against credible alternatives; test reliability, sensitivity, failure modes, and downstream consequences; document access and versioning.

**Common combinations**

- Evidence Synthesis and Qualitative Methods for need and construct discovery;
- Experimental or Computational Research for validation;
- Reproduction & Replication for independent usability and robustness.

**Exit**

- communicate the artifact together with its validation boundary, failure modes, version, and evidence—not as a standalone implementation claim.

## 9. Reproduction & Replication

**Use when**

- checking whether reported computations can be regenerated, whether a result recurs under the same or meaningfully varied conditions, or whether an influential claim is robust;
- validating an experiment, benchmark, dataset, analysis, or method implementation.

**Starting conditions**

- the target claim, source, protocol, code/data availability, expected outputs, and acceptable correspondence criteria;
- a declared type: computational reproduction, direct replication, conceptual replication, robustness test, or independent reimplementation.

**Produces**

- evidence about artifact correspondence, result repeatability, robustness, transport, or implementation dependence;
- a discrepancy diagnosis separating source ambiguity, environment drift, implementation error, sampling variation, and scientific contradiction.

**Cannot establish alone**

- fraud or incompetence from a failed attempt;
- universal truth from a successful repetition;
- direct replication when material protocol details changed.

**Validity risks and quality floor**

- silent deviations, incompatible environments, underpowered replication, outcome switching, asymmetric troubleshooting, and calling a conceptual test “direct”;
- predefine correspondence and analysis; record deviations and assistance; compare effect estimates and uncertainty, not only significance labels; preserve successful and failed attempts.

**Common combinations**

- Computational or Experimental Research for execution;
- Evidence Synthesis to interpret the replication landscape;
- Method & Artifact Research when the target is an instrument, dataset, or benchmark.

**Exit**

- classify reproduction or replication as consistent, inconsistent, mixed, inconclusive, or method-failed within the attempted correspondence. Update the target claim without erasing the original record.

## Composition patterns

These are defaults, not fixed pipelines:

| Goal | Typical minimum composition |
|---|---|
| Propose and validate an innovative claim | Evidence Synthesis + Theory Building + the most discriminating empirical mode + challenge/replication as warranted |
| Assess whether an idea is worth pursuing | Evidence Synthesis + Exploratory Discovery or Theory Building; stop at positioning |
| Validate an existing experiment | Experimental or Computational Research + Reproduction & Replication + bounded inference |
| Write an original empirical paper | Relevant modes across the inquiry lifecycle; do not add unrelated modes merely because the output is a paper |
| Produce a literature review | Evidence Synthesis, optionally Theory Building or meta-analysis; no default experiment |
| Investigate an unclear phenomenon | Exploratory Discovery + an appropriate observational, qualitative, or computational mode, followed by independent confirmation if claims advance |
| Build a benchmark or scientific tool | Method & Artifact Research + relevant validation mode + replication |

If a selected method cannot distinguish the central claim from its strongest rival, revise the method combination before collecting more evidence.
