Skip to content

On-device language

TinkyMind

Local language. A voice that stays yours.

Model profile · 8 October 2026

Introducing TinkyMind, our small local language model for communication. It proposes replies on the device, without a cloud connection. The surrounding TinkySpeak runtime keeps proposals, the user’s choice, and confirmed delivery separate.

The documented checkpoint is child-trained and English-only. TinkyMind is a neural language model with weights; local execution is its deployment choice, not a claim that every suggestion is correct.

Capabilities

What TinkyMind is built for

01

Keep inference on the device

Run the installed ONNX pack on the CPU in a bounded local worker.

02

Respect the checkpoint’s scope

Declare language and audience limits instead of borrowing another model’s coverage.

03

Let the user choose

Suggest possible replies without automatically making them the user’s delivered words.

Explore a decision

Situation

A conversation partner asks what you would like.

Choices, not speech
  1. Run the local checkpoint
  2. Return possible reply choices
  3. Wait for the user rather than deliver automatically

Protocol illustration—not a live TinkyMind generation.

The evidence so far

documented model-pack size
24 MB
local inference
CPU
current checkpoint
English

The offline documentation records an installed-CLI check under a network-denying rule. The current report also inspects manifest validation, audience and language refusals, and worker limits. The pack size and historical offline check were not remeasured for that report. See methods and limitations →

Compute & cost

A different compute path

The local pack avoids a cloud model call for its inference path. Device memory, inference time, battery use, and integration still matter: a 24 MB artifact is not a measurement of peak runtime memory. Evaluate the exact checkpoint on the device you plan to deploy.

Boundaries

Know where the guarantees begin—and end

A suggestion is not the user’s voice until they choose or compose it and request delivery. The host then confirms what was delivered. This boundary preserves control, but does not establish clinical efficacy or make inappropriate model suggestions impossible.

Inspect the technical account →

Getting started

Explore TinkyMind

Use the communication demo to explore the interaction. To run TinkyMind itself offline, install the local pack and dependencies described in the setup guide, then select the local provider. The hosted demo is not proof of offline execution.

Research & documentation

Go deeper into the work

Architecture, evaluation, sources, and known limitations. The complete technical account is preserved below.

Read the TinkyMind technical report

Perslis model research / TinkyMind

TinkyMind: Bounded On-Device Language for Communication

Small, local language inference for communication—without turning a proposed response into somebody’s voice before they choose it.

Model
TinkyMind
Category
On-device language
Revision
1.0 · 2026-10-08

Technical research report · research prototype · not presented as a peer-reviewed journal publication

Abstract

TinkyMind is Perslis’s local language-model option for the TinkySpeak communication runtime. The documented checkpoint is a child-trained, English-only ONNX pack, described as 24 MB, executed on the CPU in a bounded worker. Unlike Lois’s symbolic core, this checkpoint has neural weights: it proposes possible replies. The surrounding conversation protocol separately manages selection, delivery requests, and host confirmation. This report distinguishes local model execution, protocol behavior, and communication quality, none of which can stand in for the others.

1. Research question and model identity

Can a communication device propose useful replies locally while keeping the user’s words under the user’s control? The system must be usable without a cloud connection, but locality alone does not prove response quality or safety.

TinkyMind is the language model. TinkySpeak is the conversation runtime. The model’s proposed choices are not delivered speech; selection and delivery are separate protocol steps. Neither component should be renamed Lois, Peel, or FailFirst.

2. Local inference and bounded requests

  1. Partner input
  2. Local reply proposals
  3. User selects or composes
  4. Host confirms delivery

The inspected adapter identifies the checkpoint as child-v6cube-aug, requires a manifest with engine cube-onnx-v1, and checks decoding limits before generation. It launches a Python worker without a shell, sends a bounded request, validates returned choices, and stops active workers on cancellation or shutdown.

The manifest contract fixes maxPairs = 2, maxSeqLen = 256, and maxNewTokens = 20. Those are request/decoding limits, not a model-parameter count or a quality score. The documented CPU dependencies are NumPy and ONNX Runtime.

3. Language, audience, and privacy scope

The current pack declares English input and output, including supported locale variants; it does not provide translation. The adapter refuses unsupported language pairs and refuses an adult profile for this child checkpoint. Changing settings does not retrain the model or broaden its evaluated audience.

The adapter writes no transcript and its workers own no durable conversation data. That is narrower than a claim that every application or selected provider retains nothing. The host owns the conversation session and delivery path; its storage and privacy policy still need review.

Offline execution refers to an installed local pack and its runtime dependencies. It is not a claim that a hosted webpage or every TinkySpeak feature works offline.

4. Evidence and reproducibility

The existing offline documentation describes testing the installed CLI under a macOS network-denying rule, including generation, selection, delivery, history, reset, and refusal of cloud or unsupported-language switches:

node scripts/check-tinkymind.mjs

That is a documented historical runtime check. It was not rerun for this report, and the cited 24 MB artifact size was not remeasured here. Source inspection on 8 October 2026 confirmed manifest validation, language/audience refusals, bounded worker handling, and the separation of proposals from the surrounding protocol.

The separate AAC white paper reports protocol evidence and current failures. Protocol correctness is not a clinical efficacy result or proof that generated choices are appropriate for each user.

5. Limitations and next experiments

  • The checkpoint is child-trained and English-only; it is not validated for every age, language, diagnosis, or communication style.
  • It is a neural language model. Offline inference does not remove the possibility of an inappropriate or unsupported suggestion.
  • Artifact size is not peak runtime memory; CPU inference is not a measured battery-life claim.
  • User approval and host-confirmed delivery are necessary boundaries, not a universal guarantee against harm.
  • No $50-chip, 90% savings, or clinical-certification result is established here.

Further work should measure appropriate-choice coverage, unwanted suggestions, refusal usefulness, user correction burden, accessible selection, and delivery failures with AAC users and specialists. Hardware evaluation must report the checkpoint digest, device, peak memory, latency distribution, and measured energy under an offline rule.

6. Primary sources and related research

  1. TinkyMind models and privacy documentation: documented pack, CPU execution, offline checks, and checkpoint scope.
  2. Perslis.AAC: User-Controlled Conversation: protocol, proposal-to-delivery boundary, evidence status, and failures.
  3. Inspected implementation: api/_lib/aac_runtime/tinkymind.js and capabilities.js, local checkout reviewed 8 October 2026.
  4. Communication-runtime commands: model selection, reset, and explicit delivery.

The Perslis model family

Different models. Different jobs.