One memory ceiling / Apple silicon

Run the model.
Not the memory bill.

You set one memory ceiling. Astronomical is a local model runner for language, vision, and image models: sparse experts move between memory and solid-state storage when they do not fit. Every request stays on this Mac.

ACTIVE GPU selected experts

SSD / cold experts

MLX / unified memory

Product facts
not benchmark claims
1user memory ceiling
3chat, vision, images
0cloud accounts
LoopbackOpenAI-compatible API

01 / On this Mac

Install. Open Library. Point a local client.

Astronomical is a local model runner for one user on Apple silicon. You set one memory ceiling; sparse experts stream between unified memory and solid-state storage when the full payload does not fit. The native menu and Observatory console show that state. Stable requires macOS 14 (Sonoma) or later, declared best-effort across the M-series spectrum.

Menu See the ceiling while you work

The menu reports effective memory, expert residency, graphics-processor use, prompt progress, and prompt reuse.

Observatory The local console

Observatory lives on the same loopback address. Open Library, watch status, and send a chat without a separate monitoring stack.

Local API Use the tools you already have

Chat Completions, Responses, and image generation speak OpenAI-compatible JSON over loopback only. Astronomical loads the requested model on demand.

02 / Capabilities

What the app can do today.

These are shipped product surfaces. Speed still depends on the model, context, storage, and the memory you leave for the rest of the Mac.

01 Chat locally

Send a conversation through Observatory or any local OpenAI-compatible client. Tool calls and Responses share that same loopback server.

02 Vision in the same chat

Attach an image to a request. A vision-capable model sees it on this Mac. The image does not leave the laptop.

03 Generate an image

A text prompt becomes one lossless PNG through the local image endpoint. Text and image models share the app, not a second service.

04 Install from Library

Observatory can download a curated public model onto this Mac. Astronomical verifies the payload before the model becomes discoverable.

05 One memory ceiling

Set the model RAM your laptop can spare. Astronomical keeps complete experts resident when they fit and streams selected experts from storage when the request needs room.

06 No hidden quantization

A tighter ceiling does not silently downgrade the model. Paging keeps the artifact's declared types. Unsupported files fail validation instead of being guessed from a name.

03 / Automatic expert residency

Fully in memory when it fits.
Paged when the request needs room.

When capacity allows Fully in memory

Complete sparse-expert layers stay resident for the lowest-latency path.

When the request needs room RAM + SSD streaming

Complete layers make room and Astronomical pages only router-selected experts from solid-state storage.

One memory ceiling. No manual mode switch. The same loaded model adapts to live context and runtime memory needs, then restores complete expert layers when capacity returns.

04 / Field reports

Claims need receipts.

Each report labels architecture facts, acceptance evidence, and unresolved questions separately.

R–001 Verified

Prompt reuse / Persistent state

A repeated prefix skipped 2,048 prompt tokens.

The first request populated one state block. The identical second request restored it while preserving generated-token parity.

Read report R–001
NEXT Planned

Representative measurement

One thousand in. One thousand out.

The next public comparison will use a representative prompt and generation workload, report memory alongside speed, and retain failed hypotheses.

See the evidence standard

05 / Evidence standard

Fast is a measurement,
not an adjective.

  1. 01
    Representative work

    Performance claims use roughly 1,000 input and 1,000 output tokens, with the exact workload disclosed.

  2. 02
    Memory beside speed

    Throughput is reported with model precision, effective memory ceiling, and peak Machine Learning framework memory.

  3. 03
    Cold and warm separated

    Operating-system file reuse is not presented as physical solid-state-drive performance.

  4. 04
    Failed ideas stay visible

    A smaller command count or prettier trace does not become a win unless end-to-end behavior improves.

Built in the open

The laptop is the deployment target.

No cloud account. No model upload. No hidden precision downgrade.