The menu reports effective memory, expert residency, graphics-processor use, prompt progress, and prompt reuse.
One memory ceiling / Apple silicon
Run the model.
Not the memory bill.
You set one memory ceiling. Astronomical is a local model runner for language, vision, and image models: sparse experts move between memory and solid-state storage when they do not fit. Every request stays on this Mac.
SSD / cold experts
MLX / unified memory
not benchmark claims
01 / On this Mac
Install. Open Library. Point a local client.
Astronomical is a local model runner for one user on Apple silicon. You set one memory ceiling; sparse experts stream between unified memory and solid-state storage when the full payload does not fit. The native menu and Observatory console show that state. Stable requires macOS 14 (Sonoma) or later, declared best-effort across the M-series spectrum.
Observatory lives on the same loopback address. Open Library, watch status, and send a chat without a separate monitoring stack.
Chat Completions, Responses, and image generation speak OpenAI-compatible JSON over loopback only. Astronomical loads the requested model on demand.
02 / Capabilities
What the app can do today.
These are shipped product surfaces. Speed still depends on the model, context, storage, and the memory you leave for the rest of the Mac.
Send a conversation through Observatory or any local OpenAI-compatible client. Tool calls and Responses share that same loopback server.
Attach an image to a request. A vision-capable model sees it on this Mac. The image does not leave the laptop.
A text prompt becomes one lossless PNG through the local image endpoint. Text and image models share the app, not a second service.
Observatory can download a curated public model onto this Mac. Astronomical verifies the payload before the model becomes discoverable.
Set the model RAM your laptop can spare. Astronomical keeps complete experts resident when they fit and streams selected experts from storage when the request needs room.
A tighter ceiling does not silently downgrade the model. Paging keeps the artifact's declared types. Unsupported files fail validation instead of being guessed from a name.
03 / Automatic expert residency
Fully in memory when it fits.
Paged when the request needs room.
Complete sparse-expert layers stay resident for the lowest-latency path.
Complete layers make room and Astronomical pages only router-selected experts from solid-state storage.
One memory ceiling. No manual mode switch. The same loaded model adapts to live context and runtime memory needs, then restores complete expert layers when capacity returns.
04 / Field reports
Claims need receipts.
Each report labels architecture facts, acceptance evidence, and unresolved questions separately.
Prompt reuse / Persistent state
A repeated prefix skipped 2,048 prompt tokens.
The first request populated one state block. The identical second request restored it while preserving generated-token parity.
Read report R–001Representative measurement
One thousand in. One thousand out.
The next public comparison will use a representative prompt and generation workload, report memory alongside speed, and retain failed hypotheses.
See the evidence standard05 / Evidence standard
Fast is a measurement,
not an adjective.
- 01Representative work
Performance claims use roughly 1,000 input and 1,000 output tokens, with the exact workload disclosed.
- 02Memory beside speed
Throughput is reported with model precision, effective memory ceiling, and peak Machine Learning framework memory.
- 03Cold and warm separated
Operating-system file reuse is not presented as physical solid-state-drive performance.
- 04Failed ideas stay visible
A smaller command count or prettier trace does not become a win unless end-to-end behavior improves.
Built in the open
The laptop is the deployment target.
No cloud account. No model upload. No hidden precision downgrade.