Before this project, I understood those pieces separately. Making them depend on one another changed what I paid attention to. A record identifier is not just two integers. A page is not just a byte array. Even a projection operator needs a precise rule for who owns the returned tuple and how long its values remain valid.
I started where the abstractions stop
PageDB begins with fixed-size persistent pages managed by a disk manager. Above that, slotted pages store variable-length records. The slot directory lets a record keep the same page-and-slot identity while its bytes move during compaction. That distinction became the first useful lesson: logical identity and physical position should not be the same thing unless I am willing to make every caller absorb layout changes.
The buffer pool made resource ownership harder. A page can be resident, pinned, dirty, or eligible for Clock replacement. Evicting a pinned page is wrong; forgetting to mark a changed page dirty silently loses work; retaining a pointer after unpinning lets a later eviction invalidate it. What looked like a cache was really an agreement about lifetime.
SQL request
→ parser / binder / planner
→ pull-based operators
→ catalog and typed tuples
→ table heap / B+ tree
→ buffer pool
→ slotted page
→ disk manager
Record IDs connected the storage layers
A table heap links pages and exposes records as stable RecordId values. Insertion may choose another page, and records can move within a slotted page, but callers continue to address the record through its page and slot. Multi-page scans forced me to define what happens at page boundaries and how a scan advances without leaking pins.
The current heap is intentionally simple: a linked page chain, linear first-fit insertion, and linear membership validation. Records cannot span pages. Deleted slots remain tombstones rather than being reused. Those decisions kept the format inspectable, but they also give the design concrete limits: a page can exhaust slot-directory space even when record bytes have been deleted, and insertion does more scanning than a free-space map would require.
The index could not pretend failures were atomic
PageDB's B+ tree supports unique signed 64-bit keys, exact lookup, ordered forward traversal, insertion, and rebalanced deletion. Splits and merges made persistence feel different from an in-memory data-structure exercise. A root change has to survive reopening, siblings and parents must agree after structural edits, and every referenced page must be reachable.
But there is no transaction layer or write-ahead log yet. A failure during a split, merge, or root update can leave unreachable pages or an invalid tree. Saying that plainly matters. Passing insertion and reopen tests demonstrates the implemented path; it does not manufacture crash atomicity that the storage stack does not provide.
SQL pushed ownership upward
The execution layer uses pull-based operators: sequential and integer-range index scans, typed column-literal filters, and projections. Composing filters represents logical AND. The SQL layer parses one explicit-length SELECT statement for one catalog table, binds named columns, and plans the supported predicates.
This is not broad SQL support. There are no joins, aggregates, sorting, aliases, DML, DDL, or optimizer. SQL plans currently choose sequential scans rather than discovering catalog indexes. Keeping the grammar bounded let me concentrate on error offsets, type checking, tuple ownership, and the handoff between planning and execution.
A server made the boundary observable
The loopback C server accepts exactly one request per TCP connection using a documented binary protocol. Frames have explicit lengths and bounded fields. The server opens the database, loads a caller-supplied catalog page, runs the same SQL path used by the engine, and serializes either rows or an error.
That interface exposed stale assumptions quickly. Internal structs are not a wire format. Host byte order is not a protocol. A parser error needs a stable representation after crossing a socket. It also forced an honest inventory of missing work: the server is sequential, unauthenticated, has no TLS, and binds only to 127.0.0.1.
What I actually tested
The repository tests page serialization and record mutation, buffer eviction and dirty-page persistence, multi-page heap behavior, typed tuple encoding, catalog reopen, B+ tree insertion/search/deletion/range traversal, operator composition, SQL binding and diagnostics, and requests through the server protocol. I also run the suite under AddressSanitizer and UndefinedBehaviorSanitizer configurations.
The failures I care about are boundary failures: using a record after its page lifetime ended, reopening a structure with the wrong header identity, accepting a truncated protocol frame, or letting an index point at a stale record. Those cases say more about the engine than a happy-path SELECT.
What remains unfinished
PageDB has no transactions, WAL, crash recovery, concurrent execution, page reuse, joins, or Java client. The catalog page ID must be supplied at startup because automatic catalog discovery is not implemented. The repository README once described the Java client as future work, and the current source still agrees; I am not presenting a planned language boundary as a completed one.
The project changed how I read database diagrams. The boxes are useful, but the difficult parts live in the contracts between them: identity versus location, ownership versus lifetime, dirty state versus durable state, and a successful operation versus one that survives a crash.
Related project: PageDB case study →