METHODOLOGY V1 · CURRENT MVP HYPOTHESIS

Evidence before conclusions.

MetricMartian estimates patterns from repository activity and AI-assisted delivery baselines. These ranges are directional, not timekeeping records. Use them to identify questions and compare economic patterns—not as the sole basis for compensation or employment decisions.

READ THIS FIRST · LABELED SAMPLE

What the evidence can—and cannot—support.

SUPPORTED IN THIS SAMPLE

Seven pull requests define the window.

Period
Aug 3–9, 2026
Analyzed
6 of 7 merged pull requests
Visible range
6–9 repository-visible hours
Entered cost
$2,240 · separate from the range
Confidence
Directional · one analysis missing
NOT SUPPORTED

No conclusion about the person.

  • No exact-hours or timekeeping claim
  • No inference about unseen planning, support, or mentoring
  • No guarantee of correctness or business impact
  • No automatic compensation or employment decision

The next step is to open the Flight Recorder and ask for missing context—not to turn a partial repository record into certainty.

01

UNIT OF ANALYSIS

Evidence window: one engineer, one week.

The MVP analyzes an engineer × week using merged pull requests as evidence. Every displayed result includes its period, pull-request count, author and repository coverage, confidence or provisional state, and a path to the underlying Flight Recorder.

Responsible reading

A Mission Grade describes the current evidence window, not a person. Repository data is incomplete by design.

02

EVIDENCE LAYERS

Four layers stay distinguishable.

Repository-visible effort

Commit sessions and review cycles form a directional low-to-high range. The range describes evidence visible in GitHub and is not a timekeeping record.

AI-assisted baseline

A constrained model compares the pull request with reviewed reference examples for a competent senior engineer using modern AI tools. It returns a conservative range, confidence, and exemplar labels.

Quality evidence

Available checks, approvals, relevant test evidence, and deterministic or human-confirmed repair signals contribute. Missing signals lower coverage instead of counting as failure.

Cost and business context

Entered billed hours, monthly allocation, or a weekly override stays visibly separate from the evidence range. Founder Calibration changes value weighting, never the underlying evidence.

03

QUALITY EVIDENCE

Missing is not failing.

The quality signal can use CI/check results (35%), review approval (25%), relevant test evidence (20%), and no identified repair or revert within 14 days (20%). Only available evidence enters the score; coverage reports how much of the possible signal exists.

Decisive success

A GitHub check conclusion of success.

Decisive failure

Failure, timed out, action required, or startup failure.

Missing evidence

Cancelled, skipped, neutral, stale, pending, unknown, or unavailable.

When quality coverage is below 50%, the product shows “Limited quality evidence” and reduces overall grade confidence.

04

ECONOMICS

Every economic number keeps its denominator.

Weekly cost

Hourly work uses entered billed hours when available; otherwise it uses a clearly labeled estimate. Monthly cost is divided by 4.33 unless a visible weekly override exists.

Value points

Archetype weight × complexity multiplier × Founder Calibration. Calibration changes business-value weighting, not the baseline, effort evidence, or quality signal.

Delivery efficiency

AI-assisted baseline midpoint ÷ repository-visible midpoint. It compares two directional ranges; it does not claim how fast a person literally worked.

Cost per value point

Weekly cost ÷ total value points. When there is no delivered value or no usable denominator, the metric stays empty and the result stays provisional.

Each component is normalized against the organization's trailing eight-week median using only earlier non-provisional weeks with a usable value. The first two weeks use the visible fixed anchors; null or zero medians fall back instead of producing infinity.

05

MISSION GRADE

A transparent composite, with guardrails.

Delivery efficiency (45%), cost per value point (35%), and quality (20%) are normalized against prior non-provisional organization medians. The first two weeks use visible fixed anchors. These weights are an MVP hypothesis, not scientific truth.

0.45 × delivery+0.35 × cost+0.20 × quality
A1.15 or above

Strong current pattern

B0.95–1.14

Aligned

C0.70–0.94

Needs context

D0.50–0.69

Material variance

FBelow 0.50

Critical variance

Provisional guardrails

Week one, low-volume windows, zero-value windows, or missing quality evidence remain provisional. The one-letter grade clamp prevents one pull request from moving the displayed result by more than one letter from a prior non-provisional week; the raw composite remains available for audit.

06

CONFIDENCE AND MISSING CONTEXT

Confidence and missing context stay attached.

Overall confidence combines evidence coverage, baseline confidence, and quality confidence. A thin or conflicting signal becomes a provisional state or a clear limitation—not a more certain claim.

  • Unseen planning, support, mentoring, incidents, and collaboration are not inferred.
  • Repository-visible time and entered billed or allocated time remain separate.
  • AI output is schema-validated, bounded, and stored with its actual model and prompt versions.
  • Raw customer diffs never become global reference examples without explicit written consent.
Example: incomplete evidence stays incomplete

Six of eight pull requests have complete analysis, and three have no decisive CI result. The brief reports both counts, lowers confidence, and asks whether discovery, support, incidents, mentoring, or unpushed work explains missing context. It does not turn the gap into a timekeeping or overbilling claim.

07

VERSION POLICY

Versioned by design.

Model, prompt, deterministic narrative, embedding, scoring, reference-corpus, or methodology changes require a new version, fixture evaluation, documented deltas, and preserved historical provenance. Current evidence semantics use metricmartian-v2; the formulas and weights remain methodology v1.

Methodologymetricmartian-methodology-v1
Scoring and evidence semanticsmetricmartian-v2
Reviewed reference generationmetricmartian-reference-set-v1
Classifier promptclassifier-v16
Baseline promptbaseline-v15
Weekly narrativeweekly-notes-v2
Commit timestamp policygithub-authored-preferred-v1

Methodology change history

  1. Added the governed reference ledger and exact live-corpus verification without changing the accepted v1 corpus digest.

  2. Advanced weekly narrative evidence coverage to weekly-notes-v2 so analyzed pull requests stay separate from the total merged count.

  3. Advanced evidence semantics to metricmartian-v2 for decisive check conclusions and commit-timestamp provenance; formulas and weights did not change.

  4. Founder-reviewed and locked the first 20 synthetic reference bands as metricmartian-reference-set-v1.

Founder Calibration

A visible 0.25–3× multiplier can adjust business-value weighting for a pull request. The original value and the calibrated value are both stored. Calibration never rewrites the baseline, effort evidence, or quality signal.

08

LIMITATIONS

Limitations: what MetricMartian cannot know.

GitHub cannot prove software correctness, exact time, business impact, individual contribution inside a collaboration, or the full value of engineering work. The AI-assisted baseline can also be wrong. MetricMartian keeps these limitations visible and recommends a question—not an employment action.

Responsible-use boundary

Use MetricMartian to identify questions and compare economic patterns. Do not use a Mission Grade as the sole basis for compensation, discipline, hiring, termination, or any other employment decision.