Research and publications

This page is the canonical record of my research. My work spans quantitative finance, financial machine learning, large language models, information extraction and mathematical modelling, with a focus on empirical and reproducible research at the intersection of markets and machine learning. For current engineering work, model releases and smaller reproducible projects, see GitHub.

FinVector-Market-4B: Measuring Domain Fine-Tuning Gains Beyond Prompt Compliance in Financial Reasoning

Alina Khaybullina. 2026.

Abstract. This study investigates how much domain-specific supervised fine-tuning improves a small language model for financial reasoning once improvements attributable to prompt structure are separated from improvements attributable to training. FinVector-Market-4B adapts Qwen3.5-4B using low-rank adaptation on a curated financial-reasoning corpus and evaluates the resulting model across structured macro interpretation, financial question answering, calculator routing, causal transmission and multi-asset scenario-analysis tasks.

The evaluation uses a frozen 600-example benchmark and controlled decoding conditions to compare the base and adapted models under the same explicit output contract. Domain adaptation produces substantial gains across several task families, including financial question answering, calculator-expression generation and structural scenario reasoning, while also materially improving structured-output reliability.

The paper focuses on a narrower and more useful question than whether fine-tuning simply makes a financial language model “better”: which capabilities actually improve after adaptation, once prompt engineering is controlled for? The results suggest that targeted domain adaptation can produce genuine gains in financial task competence, while also revealing where those gains fail to generalise.

Machine Learning and Dynamic Risk Allocation in Trend Following

Alina Khaybullina. 2026. Preprint.

Abstract. This study examines whether a nonlinear machine-learning overlay can improve the risk efficiency of a transparent trend-following strategy across eight international equity ETFs. A conventional twelve-month trend and volatility rule first determines each fund’s baseline allocation. A histogram gradient-boosted classifier then makes one decision: retain an eligible sleeve when estimated underperformance risk is acceptable, and hold cash otherwise. A bounded set of model and allocation configurations is ranked on 2010–2015 after-cost certainty equivalent. The selected combination is evaluated from January 2016 through December 2025 using annual chronological refits, fully matured labels, next-close execution and a three-basis-point one-way trading cost. The overlay produced a stronger historical return-risk profile than the baseline, but a wide bootstrap interval and a later-period return shortfall limit any claim of persistence.

Cross-Asset Shock Diffusion: A Reproducible Test of Residual Underreaction, Shock Coherence, and Trading Economics

Alina Khaybullina. 2026. SocArXiv preprint.

Abstract. This paper examines whether differences in the speed with which traded assets respond to a common market shock can predict subsequent relative returns. The framework combines a lagged rolling factor model with Absorption Gap, which measures an asset’s response error, and Shock Coherence, which characterises the market state on that date. The study evaluates the signal using executable next-open timing, explicit transaction costs, dependence-aware inference, randomised-signal benchmarks, chronological diagnostics, alternative normalisations, portfolio sensitivity analysis and machine-learning extensions. The sample contains 24 ETFs from 4 January 2010 through 28 August 2026, with eight factor proxies excluded from the 16-asset traded cross-section. The contribution is both empirical and methodological: it develops a transparent framework for measuring relative shock absorption, separates cross-sectional signal information from market-state effects, and emphasises timing integrity, transaction costs, robustness, falsification, and reproducibility.

NLP Methods for Information Extraction from Text

Alina Khaybullina. 2022. Technical manuscript.

Abstract. The manuscript surveys how natural-language processing converts unstructured text into usable information, with emphasis on word and sequence analysis. It covers speech recognition, probabilistic language models, sequence labelling, vector semantics, word embeddings and text classification, and includes a sentiment-classification exercise using more than 10,000 user reviews.


ORCID: 0009-0007-2586-842X.

Last reviewed: 19 September 2026.