Scientific Method Engine

LEARN's operational research framework for rigorous investigation of intelligence as a measurable physical phenomenon.

The Scientific Method Engine transforms LEARN from a documentation system into an experimental research instrument by implementing an 11-step operational cycle, explicit falsification criteria, controlled experiments, and provenance tracking for every modification.

The Operational Research Cycle

11-step process from observation to re-observation

1

Observe

Capture empirical information from reality or experimental environment

"What exists?"

2

Record

Timestamp and preserve observations with full provenance

"What was measured?"

3

Represent

Transform observations into structured knowledge and models

"What does this mean?"

4

Hypothesize

Generate competing theoretical explanations

"What explains this?"

5

Predict

Preregister testable predictions before experimentation

"What comes next?"

6

Experiment

Perform controlled tests with clear independent and dependent variables

"What changes?"

7

Measure

Quantify outcomes and compute residuals

"What is the discrepancy?"

8

Falsify/Support

Compare results against predictions and falsification criteria

"Is the hypothesis supported?"

9

Adapt

Modify architecture, models, or strategy only when evidence warrants

"What changes?"

10

Verify

Independently confirm changes produce predicted effects

"Is the adaptation valid?"

11

Re-observe

Return to Step 1 with updated models and residuals

"What is the new state?"

The cycle returns to Step 1 with updated representations and residual measurements.

Never skip steps. Record incomplete cycles.

Critical Epistemological Distinctions

Never collapse these categories—each represents a different relationship to reality

Observed

Directly measured or recorded from sensors/experiments

Example: Website interaction timestamp, pixel values from camera

Inferred

Derived from observations through logical or statistical processing

Example: User intent inferred from interaction patterns

Predicted

Generated before observation, preregistered

Example: Expected user behavior given hypothesis H-001

Experimentally Tested

Prediction was measured against controlled reality

Example: Hypothesis H-001 tested via A/B architecture change

Adapted

System or model modified based on evidence

Example: Memory service revised after predictive residual decreased

Externally Validated

Result confirmed by independent observer/system

Example: Independent lab replicated prediction accuracy

Principle: A system cannot establish the validity of its own intelligence claims solely through its own evaluation. Internal coherence ≠ empirical evidence. Logical residual reduction ≠ predictive residual reduction. Website improvement ≠ intelligence improvement.

Scientific Status Labels

Every claim must be explicitly classified

Conceptual

Theoretical framework, not yet operationalized for testing

Operationalized

Defined measurable variables and testing methodology

Experimentally Tested

Hypothesis tested via controlled experiment

Piloted

Initial small-scale experiment executed and documented

Replicated

Experiment repeated with consistent results

Inconclusive

Evidence does not support or clearly falsify hypothesis

Falsified

Evidence contradicts prediction or falsification criterion met

Externally Validated

Independent evaluator confirmed result

Critical Requirement

Every major claim on the LEARN website must display one of these status labels. Avoid generic language like "proven," "validated," or "solves intelligence." Use precise language: "proposes," "hypothesizes," "tests," "investigates," "measures," "explores."

What the Scientific Method Engine Enables

From documentation to experimental infrastructure

Hypothesis Testing

Every significant architectural modification is tied to an explicit hypothesis, prediction, and falsification criterion. Tests are preregistered before execution.

Controlled Experiments

Autonomous interventions are compared against human-directed, random, and no-intervention controls. Difference is measured as residual delta.

Traceable Provenance

Every modification records timestamp, hypothesis, prediction, evidence, residual delta, and agent responsible. History is never overwritten.

Falsification Priority

LEARN optimizes for experiments that could disprove its hypotheses. Falsified hypotheses are preserved as data, not deleted.

Discrete Cognitive Segmentation Events in the Research Cycle

Temporal markers of state transitions

The 11-step operational cycle naturally aligns with discrete cognitive segmentation events — automatic transition points where the system moves between processing states. These events are measurable, load-dependent temporal signatures of cognition itself.

Observation-Representation Boundary

Step 1-3 transition: Raw sensory data commits to a structural model. A segmentation event marks where the system consolidates observations into representation.

Event type: Representation consolidation

Hypothesis-Prediction Commit

Step 4-5 transition: After generating competing hypotheses, the system commits to a specific prediction. A segmentation event marks this decision boundary.

Event type: Prediction commit

Verification-Adaptation Consolidation

Step 8-9 transition: After verification, the system consolidates results and commits to model updates. A segmentation event marks the consolidation point.

Event type: Adaptation consolidation

Load-Dependent Timing

The frequency and timing of segmentation events vary with information-processing demand:

  • High uncertainty: Frequent segmentation events, rapid state transitions, short consolidation windows
  • Resolution phase: Consolidation events concentrated around hypothesis validation or adaptive decisions
  • Stable state: Segmentation event suppression, long continuous processing windows, maintenance mode

The Scientific Method Engine is the Foundation

The next layer is the Residual Observatory, which implements six categories of measurable residuals:

Logical

contradictions, undefined terms

Predictive

prediction vs. observation

Empirical

claims vs. evidence

Adaptation

modification efficacy

Embodiment

digital vs. physical

Validation

self vs. independent

Next: Residual Observatory