Research note 8 min read

A perfect braid, an ordinary classifier

Exact algebra survived every tested braid rewrite, yet lost on MNIST. A guarantee helps learning when the variation it removes is irrelevant to the task.

  • Learning & cognition
  • Applied modelling
  • Institutions & evaluation
EXACT RELATIONMNIST · MEASURED ACCURACYLearned braid model95.12%Small conventional CNN98.56%Three seeds · CPU only · 10,000 held-out digits

We built a neural-network prototype around an appealing proposition: some of the computation a model learns can be replaced by mathematics it cannot violate. The representation satisfied the braid relation by construction. Equivalent braid words produced the same features. A small classifier remained consistent even when a short word was expanded to a thousand crossings.

Then we asked it to recognize handwritten digits. A conventional image model was more accurate, smaller, and faster.

The useful finding is the gap between those two outcomes. A guarantee becomes a learning advantage when the distinctions it removes are irrelevant to the task. On our braid benchmark, they were. On our first image benchmark, we had supplied no reason that they should be.

A relation the model did not have to learn

A braid word describes crossings between strands, in sequence. Different words can describe the same element of a braid group. In the three-strand case, two adjacent crossing operations obey the relation σ₁ σ₂ σ₁ = σ₂ σ₁ σ₂. A crossing followed by its inverse also changes nothing.

A general sequence model receives these as different sequences of symbols. It can learn that some presentations are equivalent, but its architecture does not require that understanding. We used the reduced Burau representation: small matrices whose products satisfy the group relations algebraically. The neural network learned a classifier on those matrix features.

That changes the training problem. A labelled example can cover many equivalent presentations because they already reach the same input to the classifier. Training no longer has to discover that particular equivalence from examples.

In a controlled follow-up, we trained on short presentations of eight selected braid classes and tested equivalent words with up to 1,000 crossings. Across three seeds, the full matrix model achieved 100% accuracy on those long words. The GRU baseline averaged 26.04%; the small Transformer averaged 12.50%. Both baselines had received the same selected classes and training presentations.

These were known braid classes under changes of presentation. We had not demonstrated recognition of arbitrary new braid elements. Candidate classes were checked for separation in the chosen features before training; the dataset therefore tested classes those features could distinguish. The experiment established a mechanism, with a deliberately favourable alignment between the representation and the label.

We also encountered a less photogenic result. Different randomly generated words sometimes received different labels despite having indistinguishable algebraic features. Distinct strings are not necessarily distinct braid elements. We retained that diagnostic and corrected the class-selection procedure before interpreting the follow-up accuracy. A dataset can contradict its own mathematics before a model makes its first mistake.

The bridge to images was our hypothesis

The next question was whether the same computation could do something useful outside a task constructed around braid equivalence. We chose MNIST, the standard handwritten-digit dataset, as a simple external test.

Our first conversion scanned each small image patch and assigned its pixel intensities to a short braid word. Four matrix channels then produced features for a classifier. We used a change of basis that made the generators unitary: their products remained bounded, while their algebraic relations were preserved.

The matrices had mathematical justification. The pixel conversion had experimental justification only. Nothing in the definition of a handwritten three says that replacing one encoded braid word with an equivalent word should preserve the digit label. If two visually different patches reach the same representation, the classifier cannot recover the difference from those features.

The fixed conversion averaged 80.11% test accuracy. An equally sized matrix encoder without the braid relation averaged 80.59%. A small model operating directly on pixels reached 97.34%.

We then allowed a small convolutional front to learn how image patches should become discrete words. Each forward pass still used actual generator products. Training used a straight-through approximation to the gradient of the discrete choices; the optimization was approximate, while the forward algebra remained the implemented representation.

That brought the braid model to 95.12%. Learning the conversion recovered considerable performance. The next comparison determined how much of that recovery we could attribute to the braid relation.

The control that changed the story

The matched control used the same learned image front, classifier, and number of parameters. Its crossing matrices had the same eigenvalues and exact inverses, but did not satisfy the braid relation. It reached 95.57%.

The small conventional convolutional neural network, or CNN, reached 98.56%. It used roughly a quarter as many learned parameters as the learned braid model and had lower measured inference latency.

Model MNIST accuracy Learned parameters CPU ms per image
Fixed braid conversion 80.11 ± 0.31% 101,066 0.144
Learned conversion + braid matrices 95.12 ± 0.64% 102,994 0.157
Learned conversion + matrix control 95.57 ± 0.05% 102,994 0.157
Small conventional CNN 98.56 ± 0.22% 26,698 0.081

Accuracy is the mean and standard deviation across three initialization and shuffle seeds. Each model trained for ten epochs on 55,000 images, used 5,000 separate validation images for checkpoint selection, and was tested on the original 10,000 held-out images. All runs used CPU execution. Latency is the mean across seeds of median batch-one timings, including image conversion, on our local machine with two PyTorch threads. It is a prototype wall-clock measurement, not a hardware-independent performance claim.

The full measured comparison includes the pixel models and a continuous-feature CNN with nearly the same size as the learned braid model. We also reloaded all 24 saved image-model checkpoints and independently reproduced their reported test scores. Code, weights, checksums, configurations, and raw measurements are available in the experiment archive.

The matrix control's small accuracy lead does not establish a general superiority from three seeds. It does prevent us from attributing the learned encoder's improvement to the braid relation. The conventional baselines also remove any basis for claiming an MNIST accuracy, size, or inference-speed advantage for this prototype.

The guarantee survived; the advantage did not transfer

We checked whether training had broken the property that motivated the project. For the first 128 held-out images in each learned-word run, we inserted cancellation pairs and an identity derived from the braid relation into the encoded words.

The braid model preserved its predictions in every tested rewrite. Under the braid-relation identity insertion, the matched matrix control preserved an average of 84.11%. In the learned runtime's numerical precision, the braid generators' maximum relation error was approximately 9.54 × 10⁻⁸.

The model therefore had the intended structural property. That property did not translate into better ordinary digit recognition. The experiment makes a distinction that benchmark reporting can easily blur: consistency under a specified transformation and accuracy on a useful task are separate measurements.

This connects to our argument for metrics that still discriminate. The persuasive endpoint would have been “exact algebra recognizes digits at 95%.” The informative ledger includes the failed fixed conversion, the matched control, the smaller CNN, and the additional cost of the algebraic path. Those records change the decision about what to build next.

What this experiment leaves open

We have evidence about these implementations, conversions, and training budgets. We have no result showing that every algebraic image representation must lose to a CNN. The learned conversion followed inspection of the first experiment, so this is an exploratory study rather than a preregistered final comparison. Three seeds provide a useful check on repeatability, not a comprehensive search over architectures or hyperparameters. We did not measure energy consumption.

The braid experiment also remains separate from full knot recognition. Trace and determinant features gave us conjugation invariance, a useful step toward reasoning about closed braids. We did not implement a closure representation guaranteed to survive changes in strand count under Markov stabilization. A general knot-classification claim would therefore exceed the work.

Three questions would justify revisiting the method. Can a real task supply braid equivalences that demonstrably preserve its labels? Does enforcing those equivalences reduce the data or compute required against a strong conventional baseline? Can a learned encoding preserve the information that the task needs while making the mathematical guarantee useful?

Until those questions have a concrete application, a further increase in MNIST accuracy would be a weak reason to continue this particular architecture. We can retain the implementation and the negative result without assigning them an indefinite research budget.

A structural guarantee earns its place in a learning system when it removes variation the task can afford to forget. The proof establishes the guarantee; the benchmark must establish its value.

For research partners whose mandate includes reliable models under known equivalences, we welcome a concrete problem and a testable operating budget. The starting point should be an application with a reason to need the mathematics.