Files
PgTidy/plan.md
T
2026-06-23 16:59:50 +02:00

145 lines
7.9 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# PgTidy — PostgreSQL Linter & Formatter
## Context
We are starting a greenfield project, **PgTidy**: a PostgreSQL-focused linter and
formatter that enforces consistent SQL style, detects PostgreSQL-specific issues, and
formats migrations, schemas, functions, procedures, triggers, and related DB code.
The primary day-one use case is **formatting PL/pgSQL stored procedures** to match an
existing house style (sample corpus at `testdata/corpus`).
The core is written in **Go**. We ship a CLI, then an LSP server that powers a **VSCode
extension** (priority) and later a **DataGrip/JetBrains** integration (lower priority).
### Decisions locked in (from planning Q&A)
- **Parsing = hybrid.** Real PostgreSQL grammar via **go-pgquery** (libpg_query compiled
to WASM with `wazero`, **no cgo**) powers deep semantic lint rules. A **custom lossless
lexer + CST** powers the comment-preserving formatter. (libpg_query drops comments &
whitespace, so it cannot be the sole basis for a formatter — confirmed.)
- **V1 = Formatter + CLI first.** Lint → LSP/VSCode → DataGrip follow.
- **Lint scope (later milestones):** style/consistency, migration safety (Squawk-style),
naming conventions, correctness/anti-patterns — all four.
- **Editors:** one Go LSP core → VSCode now (bundled per-platform binary), DataGrip via
the free **LSP4IJ** plugin later.
## Why this architecture
- **No cgo** (WASM-embedded libpg_query) keeps cross-compilation trivial and lets us bundle
a single static binary per platform inside the VSCode extension. sqlc migrated to this exact
approach to escape cgo pain.
- **Formatting needs a lossless CST** (round-trippable, comments + whitespace preserved). This
is the dominant pattern in mature tools (Roslyn, rust-analyzer/rowan, Biome, SQLFluff). We
build our own lexer/CST so the formatter never loses a comment.
- **Deep lint needs an accurate AST** — go-pgquery gives the real PostgreSQL parse tree.
- **One core, many frontends** — CLI, LSP, VSCode and DataGrip all reuse the same engine and
the same `diagnostics` type.
## Proposed layout
```
pgtidy/
cmd/pgtidy/ — CLI entry; subcommands: fmt, lint (v2), lsp (v3), version
pkg/lexer/ — lossless lexer: tokens + trivia (comments/whitespace) attached
pkg/cst/ — concrete syntax tree (round-trippable node model)
pkg/parser/ — recursive-descent parser → CST (DML + DDL + PL/pgSQL)
pkg/format/ — Doc-IR printer (Wadler/Prettier-style) + style application
pkg/pgast/ — go-pgquery wrapper: SQL → real PG AST (for lint, v2)
pkg/lint/ — rule engine + rule packs (v2)
pkg/config/ — .pgtidy.yaml discovery/merge: style + rule config
pkg/diagnostics/ — shared diagnostic type (CLI + LSP)
pkg/lsp/ — LSP server (v3)
editors/vscode/ — VSCode extension (TS), bundles pgtidy binary (v3)
editors/datagrip/ — LSP4IJ integration (v4)
testdata/ — golden formatter fixtures + lint fixtures + corpus
```
## House style (formatter defaults — reverse-engineered from the corpus)
These become the default `style` config; all are configurable. The corpus had human
inconsistencies (e.g. a DECLARE var at column 0, mixed `=`/`:=`); the formatter **normalizes**
to the intended style below.
- Keywords **UPPERCASE**; data types **lowercase**; identifiers lowercase snake_case.
- Indent: **2 spaces** per level.
- **Leading-comma** style, one item per line, for SELECT column lists and function params.
- Function headers: params one-per-line in parens; `LANGUAGE` / `SECURITY` / volatility each
on own line; `AS` then `$$` on its own line; body; closing `$$;` on its own line.
- PL/pgSQL: `DECLARE` alone, vars 2-space indented; `--Block--` comment markers preserved;
`BEGIN`/`END` at body level; `IF/THEN/ELSIF/ELSE/END IF`, loops, `CASE` indent their bodies.
- Spacing: spaces around binary operators (`=`,`<>`,`||`,…) and `:=`; **no** space around
`::`, `->`, `->>`, array `[...]`, or before a call's `(`.
- Dollar-quote tags preserved verbatim (`$$`, `$S$`, `$Z$`, …).
## Milestones
### V1 — Formatter + CLI (priority)
1. **Lexer** (`pkg/lexer`): full PG token coverage incl. dollar-quoted strings, `--` and
`/* */` comments, operators. Comments + whitespace captured as **leading/trailing trivia**
on tokens. Acceptance: `emit(lex(src)) == src` byte-for-byte across the corpus.
2. **CST + parser** (`pkg/cst`, `pkg/parser`): recursive descent for DML (SELECT/INSERT/
UPDATE/DELETE/CTE), DDL (CREATE FUNCTION/PROCEDURE/TABLE/INDEX/TRIGGER, ALTER, DO).
3. **PL/pgSQL body parser**: DECLARE/BEGIN/END, IF/CASE/LOOP, assignments, nested SQL — the
crux for stored-procedure formatting.
4. **Printer** (`pkg/format`): Doc-IR (group/indent/line/softline) driven by `style` config.
**Graceful degradation** — any span the parser can't handle passes through verbatim rather
than being corrupted.
5. **CLI** (`cmd/pgtidy fmt`): `--check`, `--write`/`-w`, stdin→stdout, `--diff`; config
discovery walking up to `.pgtidy.yaml`; CI-friendly exit codes.
6. **Config** (`pkg/config`): load/merge style config; defaults = house style above.
**Safety guarantees (tested):** semantic equivalence (re-lex output, compare non-trivia token
stream to input), and idempotence (`fmt(fmt(x)) == fmt(x)`). The corpus is the
primary safety/idempotence harness; golden-file tests for targeted cases.
### V2 — Linter
- `pkg/pgast` go-pgquery (WASM) wrapper; `pkg/lint` rule engine (rule ID, severity, config).
- Rule packs: **style/consistency**, **migration safety** (ACCESS EXCLUSIVE locks, unsafe
`ALTER`/`ADD COLUMN`, non-`CONCURRENTLY` index builds, blocking constraints), **naming**
(configurable table/column/index/constraint patterns), **correctness/anti-patterns**
(`SELECT *`, missing `WHERE` on UPDATE/DELETE, deprecated syntax). Emit `diagnostics`.
- `pgtidy lint` subcommand; `--fix` for autofixable rules.
### V3 — LSP + VSCode
- `pkg/lsp`: `textDocument/formatting` + range formatting, `publishDiagnostics`, `codeAction`
quick-fixes — all reusing the core.
- `editors/vscode`: TS extension using `vscode-languageclient`, launches bundled `pgtidy lsp`.
Build **per-platform VSIX** (`win32/linux/darwin × x64/arm64`) in a CI matrix (rust-analyzer
model), with a target-less fallback.
### V4 — DataGrip
- `editors/datagrip`: integrate via **LSP4IJ** (free, works across JetBrains editions incl.
DataGrip). No core changes expected.
## Build / repo hygiene
- Replace boilerplate `AGENTS.md`/`CLAUDE.md` with PgTidy content; rewrite `Makefile`
(`APP := pgtidy`, `CMD := ./cmd/pgtidy`; keep build/test/lint/release-version targets).
- Add **goreleaser** for the multi-platform binary matrix (clean, since no cgo).
- `go.mod`: deps = `github.com/wasilibs/go-pgquery` (v2), `wazero`, a YAML lib, an LSP lib
(e.g. `go.lsp.dev/protocol`) in v3.
## Verification
- **Formatter:** `go test ./...` runs golden-file tests + the corpus harness asserting
(a) idempotence and (b) token-stream equality before/after (no semantic change).
Manual: `pgtidy fmt --diff` against several procedures; confirm output matches
the house style and comments/dollar-quote tags survive.
- **CLI:** `pgtidy fmt --check` returns non-zero on unformatted input, zero when clean.
- **Lint (v2):** fixture SQL with known violations → assert expected diagnostics; `--fix`
round-trips.
- **LSP/VSCode (v3):** load a `.sql` file in a dev-host VSCode, confirm format-on-save and
live diagnostics via the bundled binary.
## Open risks
- A lossless PL/pgSQL recursive-descent parser is the largest single effort; the
pass-through-on-unparsed fallback bounds the risk and lets us ship incrementally by construct.
- go-pgquery tracks PG17 (PG18 not yet) — fine for lint; irrelevant to the formatter path.
- Leading-comma + one-per-line is unusual vs. most formatters; it's a first-class style option,
not an afterthought.