Variational intelligence, built from the neuron up

We’re building the next generation of models: Variational Distributional LLMs.

JWAYU 杰维宇 develops language models in which uncertainty is not added only at the output. It becomes part of the internal computation itself, through EVE — the Elemental Variational Expanse.

Local distributionsinside hidden computation
Measurable uncertaintyKL, variance, mutual information
Transformer-readyintegrated into language modeling
JWAYU nine-tailed neural fox logo
EVE neuron Internal uncertainty Possible futures

Our thesis

The network is not the starting point. The neuron is.

Modern language models are extraordinarily capable, yet their internal units remain largely deterministic. JWAYU begins with a different premise: when ambiguity is intrinsic to intelligence, the computational unit should be able to represent it.

The current primitive

One input. One activation.

A conventional neuron compresses its input into a single scalar activation. It is precise, stable and efficient, but it collapses local possibilities immediately into one point.

h = φ(wᵀx + b)
The JWAYU primitive

One input. A structured space of possibilities.

An EVE unit constructs an input-conditioned posterior, samples a latent state through reparameterization and uses that state directly in computation.

qφ(z|x) = 𝒩(μ(x), diag(σ²(x)))

EVE — Elemental Variational Expanse

What is a variational distributional neuron?

It is a local probabilistic computational unit with an explicit prior, an amortized posterior and unit-level variational regularization. The activation is no longer merely a value: it is generated from a learned distribution.

1

Condition

The unit receives the current representation and produces posterior parameters.

2

Represent

A local Gaussian posterior carries a continuous set of plausible latent states.

3

Sample

Reparameterization produces a differentiable latent sample used in the forward pass.

4

Constrain

A local KL term regulates the posterior relative to its prior and makes the regime measurable.

Deterministic neuron

Point computation

Propagates one activation. Uncertainty is not represented inside the unit.

Dropout / stochastic masking

Randomized computation

Introduces randomness through masks or sampled paths, without turning each unit into an explicit input-conditioned latent posterior.

A language model should not have to collapse uncertainty before it has used it.

A variational distributional language model introduces EVE units into the model’s hidden computation, including Transformer feed-forward blocks. The overall architecture remains recognizable, while its internal computational primitive changes.

Instead of treating uncertainty as a final confidence score, the model can carry local latent structure across layers and use internal signals such as posterior variance, KL divergence and mutual information for evaluation, calibration and control.

From neuron to model

A research program across four layers.

JWAYU develops the primitive, studies its operating regime, integrates it into Transformers and turns internal uncertainty into an actionable signal.

01 · Primitive

The neuron is the distribution

Define local variational computation and make its internal probabilistic state observable.

02 · Composition

Build networks from EVE units

Study stability, capacity, latent dimensionality, temporal persistence and collapse control.

03 · Language

Variational Transformers

Integrate EVE into feed-forward computation while preserving a practical Transformer backbone.

For investors and strategic partners

A new computational primitive for uncertainty-aware AI.

JWAYU 杰维宇 is building intellectual property, experimental evidence and implementation expertise around variational distributional neurons and language models.

Why JWAYU

Change the unit, open a new model space.

1

Distinct technical thesis

Move probabilistic structure from only global mechanisms or outputs into hidden computational units.

2

Measurable internal state

Expose local signals that can support calibration, robustness analysis, routing and model control.

3

Architecture-compatible path

Develop the primitive inside familiar neural and Transformer backbones rather than replacing the entire ecosystem.

4

Published research trajectory

A growing sequence of public preprints establishes the concept, design space, language-model integration and agentic use.

Research · Engineering · Collaboration

Build the next generation of language models with us.

We welcome conversations with researchers, frontier-model teams, engineers, doctoral partners and organizations interested in uncertainty-aware multimodal and language models.