Our story

MARKETS → SYSTEMS → EMBODIED AI

The domain changed.
The mission did not.

A journey from interpreting complex environments to building systems that can act in the physical world.

01 /

Where it started

We started with a simple question:

Can intelligent systems learn useful structure from noisy, real-world environments and turn that into better decisions?

Our first proving ground was financial markets. That work became the foundation for how we approach machine learning, reinforcement learning, simulation, and intelligent control today.

We built an end-to-end quantitative research stack for ingesting and processing real-time and historical market data across multiple assets and timeframes.

On top of that infrastructure, we developed and tested a range of predictive architectures:

  • Deep Q-Networks and custom RL environments
  • Stacked LSTMs and bidirectional LSTMs
  • Multi-head attention and dueling network architectures
  • Conv1D + recurrent models and Transformer-based sequence models
  • Random Forests and automated feature selection
  • Confidence-aware probabilistic prediction

Our reinforcement-learning systems combined temporal feature extraction, experience replay, exploration strategies, custom action spaces, GPU training, and reward functions incorporating PnL, transaction costs, risk, and directional performance.

More model complexity does not automatically create more signal.

Target construction, reward design, regime dependence, data quality, execution assumptions, and out-of-sample stability often mattered more than model architecture. That realization changed the direction of the work.

02 /

From models to systems

Rather than relying on a black-box model to discover an entire strategy, we began separating the problem into distinct layers:

Data → Representation → Prediction → Decision → Execution

We built deterministic quantitative strategies and backtesting systems across NQ, ES, Gold, and their micro contracts using QuantConnect and TradingView.

The testing stack accounted for commissions and slippage, futures contract rollover, session and DST handling, position sizing, OCO execution, non-repainting signals, non-lookahead logic, and cross-market robustness.

We also integrated TradingView with AI agents through MCP, enabling programmatic chart inspection, strategy testing, multi-symbol research, and automated analysis.

Intelligence does not live inside the model alone. It lives in the entire system around it.
03 /

Where we are now

Today, we are applying those same principles to embodied AI and humanoid robotics.

Our current work focuses on integrating NVIDIA’s Generative Pre-trained Controller architecture into a Unitree G1 deployment stack.

The goal is to create a vertically integrated control pipeline spanning:

Motion representation → Task adaptation → Simulation → Inference → Physical execution
04 /

Latent motion & generative control

A major workstream involves adapting pretrained motion representations to the physical embodiment of the G1.

  • Finite scalar quantization and discrete latent motion representations
  • Joint-order canonicalization and coordinate-frame alignment
  • Action-space calibration and actuator-limit validation
  • Encoder–quantizer–decoder verification

At the policy layer, we are evaluating autoregressive latent-prior conditioning using proprioceptive state from the robot.

The architecture separates behavioral priors — reusable pretrained motion knowledge; embodiment decoding — translation into calibrated G1 actuator commands; and task adaptation — goal-directed behavior learned on top of the prior.

This allows the system to learn new behaviors without discarding the motion structure it already knows.

05 /

Reinforcement learning, revisited

RL remains part of the stack, but its role is now much more targeted.

Rather than learning locomotion from scratch, we are exploring parameter-efficient fine-tuning, supervised adapter initialization, and reinforcement-learning refinement.

The objective is to introduce task-specific behavior while constraining the policy to remain close to a valid pretrained motion distribution.

Learn the new task without forgetting how a humanoid is supposed to move.
06 /

Sim-to-real

Simulation is built around NVIDIA IsaacLab, with an independent MuJoCo path used to identify simulator-specific behavior.

We are testing observation parity, action scaling, control frequency, contact behavior, latency, sensor noise, friction variation, inertial uncertainty, and actuator response.

Domain randomization and perturbation testing are used to identify policies that work only under ideal simulated conditions.

07 /

Real-time deployment

The final layer is real-time execution. We are maintaining consistency across:

Training framework → Exported model → ONNX inference → Robot control loop

That includes observation preprocessing, history buffers, inference scheduling, output scaling, and synchronization with low-level actuation.

For robotics, timing is part of correctness.

A mathematically correct action delivered several milliseconds too late can still be physically wrong.

08 /

Instrumentation

The stack is instrumented end-to-end:

State estimation → Latent selection → Policy output → Commanded action → Measured response

This allows us to distinguish model failures from integration failures such as stale observations, frame mismatches, dropped control updates, incorrect gain application, timing jitter, or actuator saturation.

09 /

The thread that connects it all

Financial prediction and humanoid robotics may look like very different domains, but the underlying engineering problems are surprisingly similar.

Both involve noisy sequential observations, uncertainty, distribution shift, objective and reward design, simulation and validation, real-time decision-making, and strict separation of model errors from system errors.

Our early work was about building systems that could interpret complex environments and make better decisions. Today, that has evolved into building systems that can represent, adapt, simulate, and act in the physical world.

The domain changed.
The underlying mission did not.

Start a conversation

Bring us a hard
learning problem.

Discuss a research collaboration, explore a training pipeline, or request a walkthrough of our work.

Discuss your project