Observe operators, shapes, movement, and quantization.
Build intelligence in the open.
Cerebral Chips is an open-source platform for people building capable, local AI. We are making the tools, evidence, and decisions visible so a community can shape intelligence that serves the machines—and people—who depend on it.
THE MISSION
Put the tools for local intelligence in everyone’s hands.
Language models are becoming part of products, robots, industrial systems, vehicles, personal devices, and embedded interfaces. Those systems need local intelligence with predictable latency, privacy, and operation even when the cloud is unavailable.
Cerebral Chips is building that foundation as an open-source platform—from workload analysis and software integration down to the instruction architecture, memory system, and quantized datapath—so contributors can inspect it, improve it, and carry it forward together.
Turn measured behavior into explicit compute tradeoffs.
Keep intelligence responsive, private, and local.
Device measurements return to the next workload and architecture decision.
HARDWARE × SOFTWARE CO-DESIGN
Modern edge AI needs one system, not two disconnected roadmaps.
Cerebral Chips treats models, quantization, compilers, runtimes, firmware, instruction architecture, matrix and vector execution, and memory movement as one connected design problem. Real workload behavior drives the machine; hardware constraints reshape the software path.
Start with what the model actually does.
Prefill, decode, operator mix, tensor shapes, and quantization behavior establish the first design contract.
THE PROTON PLATFORM
A complete stack for on-device language intelligence.
The Proton platform connects model workloads, developer software, architecture simulation, agentic engineering, and accelerator hardware. Each layer exists to make the next layer measurable, programmable, and verifiable.
Proton LPU
A quantization-native, memory-aware language-processing accelerator architecture.
Proton SDK
A planned developer path from GGUF models and llama.cpp to runtime, kernels, and profiling.
ProtonSim
A workload and architecture simulator for choosing hardware from measurable model behavior.
Proton Forge
An evidence-gated engineering system for accelerating hardware and software iteration.
PROTON LPU / PROPOSED ARCHITECTURE
Language-model hardware designed around real inference behavior.
Proton LPU is a proposed RISC-V-controlled accelerator architecture for quantized transformer inference. It combines vector execution, quantization-aware matrix compute, explicitly managed local memory, DMA, external-memory access, device firmware, and a practical llama.cpp software path.
Logical architecture view. Interfaces and parameters remain subject to workload-driven evaluation.
Eight activation lanes interact with eight weight lanes.
GEMM, GEMV, and reduction behavior remain under evaluation.
Scale, zero-point, accumulation, and output handling stay in scope.
The first workload contract is intentionally narrow and measurable.
A 3B parameter model remains a stretch target.
Exact GGUF quantization coverage will be selected through bring-up work.
GEMM and GEMV dataflows are evaluated independently.
256 KB is a working baseline, not a finalized choice.
PROTON SDK / SOFTWARE PATH
From downloadable model to device execution.
The initial software path starts with GGUF models and llama.cpp. A Proton GGML backend, userspace runtime, firmware, kernels, command queues, and profiling tools connect model execution to the accelerator. IREE remains a future frontend and compiler integration path.
- 01GGUF modelMODEL
- 02llama.cpp / GGMLHOST SOFTWARE
- 03Proton backendHOST SOFTWARE
- 04Proton runtimeHOST SOFTWARE
- 05Firmware and kernelsDEVICE PATH
- 06Proton LPU or ProtonSimDEVICE PATH
PROTONSIM / DESIGN SPACE
Design the hardware from workload evidence.
ProtonSim models the work, memory movement, and execution behavior of real transformer workloads before architecture choices are frozen. It is intended to connect model operators and quantization formats to decisions such as vector width, matrix shape, scratchpad capacity, DMA behavior, and external-memory bandwidth.
Controls illustrate the intended design space. They do not represent benchmark results or finalized specifications.
- Selected workload
- Decode
- Quantization study
- Grouped INT4
- Compute candidate
- 8 × 8
AI-NATIVE ENGINEERING
Agents accelerate the work. Evidence decides what ships.
Proton Forge is an agentic engineering system for specification analysis, architecture exploration, RTL and verification scaffolding, firmware, kernel generation, documentation, and regression triage.
Generated artifacts are never accepted on model confidence alone. They pass through deterministic builds, reference-model comparison, randomized tests, assertions, coverage, synthesis, and human review.
KERNEL GENERATION / VERIFY
Target-specific kernels with reproducible verification.
The Proton toolchain is being designed to generate and tune custom kernels for vector and matrix targets from structured specifications. The accepted result is a versioned artifact backed by compilation, numerical differential testing, edge-case coverage, and performance evidence.
target: proton_vector_matrix
operation: quantized_linear
weights: grouped_int4
activations: int8_candidate
accumulation: int32
accept_if:
- compiles
- matches_reference
- passes_edge_casesPROTON SDK / REFERENCE ENABLEMENT
Build useful systems before custom silicon is complete.
The Proton SDK will include reference-platform enablement that reuses suitable open-source RISC-V cores, toolchains, runtimes, and ML projects to validate model execution and hardware–software integration early. These environments create testable engineering artifacts while the custom LPU architecture matures.
Reference repositories and verified bring-up notes will be published as they are ready.
A platform built in public, held in common.
Cerebral Chips is an open-source platform for everyone who wants local intelligence to be transparent, useful, and accessible. We are inviting engineers, researchers, makers, and system builders to help shape the work through ideas, review, experiments, and code.
ROADMAP / STAGE-BASED
Build the evidence, then build the machine.
The roadmap follows technical evidence rather than calendar promises. Each stage reduces uncertainty for the architecture, software path, and reference system that follows.
- 01Active
Workload contracts and RISC-V reference platforms
- 02Active
ProtonSim and architecture performance modeling
- 03Planned
Proton SDK, runtime, and llama.cpp backend
- 04Planned
Proton LPU-E0 RTL and verification
- 05Future
FPGA reference system
- 06Future
Open ASIC exploration
- 07Future
Proton LPU-E1 multi-tile edge architecture
- 08Long-term
Proton LPU-S1 server research
OPEN SOURCE / COMMON CAUSE
Build the future of local intelligence with us.
This platform is for the people who need AI to be local, understandable, and useful. Bring a question, an experiment, a review, or a line of code—every contribution can move the shared foundation forward.
