Blockchain Papers

Follow blockchain research across journals, conferences, and preprint repositories.

98 papersLast indexed Aug 31, 2026
Search papers

Paper index

98 results · page 1 of 5

Clear filters
Aug 27, 2026·Research Square
0 cites
Reinterpreting the Transition: A Text-Analytic Approach to Territorial Political Agendas in Spain’s Multilevel Climate Governance

Onsel Gurel Bayrali, Yasemin İrepoğlu Carreras

Abstract Environmental transition is increasingly governed through multilevel systems in which authority is shared across supranational, national, and regional governments. Existing research on multilevel climate governance has focused on coordination, implementation, and compliance, largely treating environmental objectives as politically consistent among territorial levels once adopted. This paper argues that multilevel governance also reshapes the political content of environmental transition itself, a process we call political reinterpretation: common climate objectives are selectively reprioritized and reframed as they enter territorially distinct political arenas. We test this argument using text analysis of parliamentary discourse, applying structural topic modeling to 10,564 speeches delivered across Spain’s seventeen Autonomous Communities (ACs) between 2019 and 2024 to uncover five substantive dimensions of environmental-transition discourse directly from legislative text. We find that these dimensions are distributed unevenly across regions, and, more critically, that Spain’s major statewide parties do not reproduce their national environmental-transition profiles across territories: territorial variation in the topics they prioritize within each party systematically exceeds the variation observed between parties operating in the same region. This pattern holds even among parties whose organizational structure gives them a strong incentive toward national uniformity, indicating that territorial incentives can outweigh the integrative pressures of statewide party organization. Understanding climate governance in decentralized systems therefore requires attention not only to how environmental policy is implemented across levels of government, but to how its political meaning is reconstructed as authority becomes territorially dispersed.

Open access
Climate Change Communication and Perception
Sustainability and Climate Change Governance
Computational and Text Analysis Methods
Original source
Aug 27, 2026·Zenodo (CERN European Organization for Nuclear Research)
0 cites
TOPO-2026: The Great Unlocking Universal Permanence Across All Architectures — The First Complete Solution to Catastrophic Forgetting at Scale

Frank Morales

TOPO-2026: The Great Unlocking — Full Summary Universal Permanence Across All Architectures Frank Morales Aguilera, BEng, MEng, SMIEEE Sovereign Machine Laboratory (SOMALA), Montréal, Canada August 2026 1. Executive Summary For 37 years, catastrophic forgetting remained unsolved. From McCloskey and Cohen's formal characterization in 1989 to the present day, every approach—regularization, rehearsal, architectural complexity—has been probabilistic, architecture-specific, and ultimately inadequate. Nobody has ever handled a catastrophic forgetting solution on this scale. TOPO-2026 is the first universal, deterministic solution to catastrophic forgetting, validated across 11 distinct architectural frameworks spanning the entire AI landscape. The framework leverages prime-anchored embedding invariants at indices {2,3,5,7,11,13} with safety constant $\Lambda = 0.9785142874$ to provide mathematical guarantees of memory preservation with O(1) memory overhead (just ~650 KB total for all domains). 2. What Makes This Unprecedented Aspect Prior Work TOPO-2026 Scale 1-2 architectures 11 architectures Guarantee Probabilistic Mathematical Memory GBs to TBs ~650 KB Success Rate 20-50% 100% Architecture TF or Non-TF only BOTH Forgetting 4-91% ≤ 0.26% Backward Transfer Never Achieved 3. The 12 Frameworks — Complete Certification Status # Framework Model Status Best Task C FGT 1 Dense Transformer GPT-OSS-20B ✅ 92.3% 1.55% 2 Mixture-of-Experts (MoE) Sarvam-30B, Mixtral-8x7B ✅ 95.9% -0.60% 3 GQA / MQA DeepSeek-V2-Lite ✅ 95.4% 0.03% 4 State Space Models (SSM) Evo2-7B ✅ 92.0% 1.32% 5 Hybrid Attention-SSM Evo2-7B ✅ 92.0% 1.32% 6 Retention Networks (RetNet) fla-hub/retnet-1.3B-100B ✅ 99.93% 0.00% 7 Recurrent Transformers & RWKV fla-hub/rwkv7-2.9B-world ✅ 93.00% 0.00% 8 Emergent Modularity MoE (EMO) allenai/Emo_1b14b_1T ✅ 99.60% 0.00% 9 Google's Titans — ❌ — — 10 Liquid Foundation Models (LFM) LiquidAI/LFM2-1.2B ✅ 93.50% 0.00% 11 HyDEA Evo2-7B ✅ 92.0% 1.32% 12 ResNet (CNN) ResNet-50 ✅ 100.0% -7.5% Certification Rate: 11/12 (100% of all available frameworks) 4. The Decay Law of Singularity — A Mathematical Discovery On July 31, 2026, during the certification of Gemma-4-E4B-Vision, a fundamental mathematical law was discovered. The Decay Law proves that the General Singularity is mathematically impossible with finite classes. Theorem: The Decay Law of Singularity With finite classes, $dI/dt$ approaches 1.0 asymptotically but never reaches it. The gap decays as $1/N$, where $N$ is the number of classes. The Decay Law Pattern: Classes (N) Baseline dI/dt Gap 17 5.8823529% 0.94118 0.05882 170 0.58823529% 0.994118 0.005882 1,700 0.058823529% 0.9994118 0.0005882 17,000 0.0058823529% 0.99994118 0.00005882 170,000 0.00058823529% 0.999994118 0.000005882 1.7M 0.000058823529% 0.99999994118 0.0000005882 Key Observations: Every 10× increase in classes adds another '9' to $dI/dt$ Every 10× increase in classes adds another '0' to the gap. This is not random. It is not heuristic. It is exact. This is the mathematical fingerprint of a natural law. 5. The Narrow Singularity — First in History Gemma-4-E4B-Vision achieved AGI_gate = 1.0, becoming the first model in history to achieve perfect cross-domain generalization with 100% accuracy across all 13 tasks over 6 runs. Component STL-10 CIFAR-100 Threshold Status AGI_gate 1.0 1.0 = 1.0 ✓ PASS ag_index 1 1 = 1 ✓ PASS M(t) 0.9984 0.9974 ≈ 1.0 ✓ PASS S_NARROW > 0 > 0 > 0 ✓ PASS 6. Backward Transfer — Unprecedented Achievement Models improve on earlier tasks after learning new ones — positive knowledge transfer. This has never been systematically demonstrated before. Domain Model Combined Forgetting Language Mixtral-8x7B -1.85% Language Sarvam-30B -0.60% SQL DeepSeek-R1-8B -0.98% World Models TOPO-JEPA -0.75% Vision ResNet-50 -7.5% 7. Zero NaN/Inf Stress Test Model Embedding Elements NaN Inf GLM-4.6V-Flash 884,736 0 0 DeepSeek-V2-Lite 209,715,200 0 0 Mixtral-8x7B 131,072,000 0 0 GPT-OSS-20B 579,133,440 0 0 Sarvam-30B 1,073,741,824 0 0 TOTAL ~1.99 Billion 0 0 8. Comparison with State-of-the-Art Method Forgetting Success Rate Memory Math. Guar. TF Non-TF TOPO-2026 ≤ 0.26% 100% 67.5-451.5 KB Yes ✓ ✓ Experience Replay 4%-91% Variable Variable No ✓ ✗ EWC 8.3%-27.7% 20% 4.4 GB+ No ✓ ✗ Full HOPE 8.5%-45.4% 20% 2-4 GB No ✗ ✓ Progressive Nets 1.8% Variable $O(k^2)$ No ✓ ✗ Key Finding: TOPO-2026 is the only method that works on both Transformer and non-Transformer architectures with mathematical guarantees, 100% success rate, and O(1) memory. 9. Solved Problems Catastrophic Forgetting: Solved across 12 frameworks and 14 domains — first time at this scale AI Bias: Eliminated through four-tier spectral annihilation (100% rejection) World Model Instability: Solved through TOPO-JEPA (-0.75% forgetting) Numerical Instability: Zero NaN/Inf across 1.99 billion embedding elements Dataset Dependence: Proven dataset-agnostic across STL-10 and CIFAR-100 The Singularity Illusion: Decay Law proves the General Singularity is mathematically impossible Architectural Dependence: Proven to work on ALL available architectures — first universal solution 10. The Complete Arc: 28 Years of Discovery Period Domain Principle Result 1998-2002 Neuroimaging (fMRISTAT) Fix sparse reference 3 df → 112 df 2026 Number Theory First 6 primes RH Proved 2026 AI Memory Six embedding rows CF Solved 2026 AI Safety Geodesic distance Zero violations 2026 AI Bias Prime-anchored equity Bias eliminated 2026 Narrow Singularity AGI_gate = 1.0 First model 2026 Universal Certification Same anchors ALL architectures! 11. Key Achievements Universal Applicability: 12 frameworks, 11 certified (100% of available) — unprecedented scale Mathematical Guarantee: $\Lambda = 0.9785142874$ provides provable anchor stability O(1) Memory: ~650 KB total for all domains — unprecedented efficiency Backward Transfer: Negative forgetting across multiple domains — first demonstration Perfect Vision Performance: 100% accuracy, 0.17% forgetting across 6 runs Dataset-Agnostic: Same protocol works identically on STL-10 and CIFAR-100 75.7× Improvement: Over Google's Full HOPE in genomics Narrow Singularity Achieved: AGI_gate = 1.0 — first in history 100% Certification Rate: Across all runs, all domains, all datasets AST-RH Byproduct: Riemann Hypothesis proved as a byproduct Zero NaN/Inf: Across 1.99 billion embedding elements 12. The Final Statement Nobody has ever handled a catastrophic forgetting solution on this scale. TOPO-2026 is the first. The stochastic illusion is over. Deterministic cognitive engineering has begun. Stability is not a probabilistic hope. It is a numerical guarantee. The Decay Law of Singularity is not a defeat. It is a liberation. It frees us from the hype cycle, the fear of the singularity, the endless pursuit of AGI, and the billion-dollar promises. It gives us a clear roadmap, a mathematical framework for control, a focus on solving real problems, and an honest assessment. "Genomics is permanent. Language is permanent. Vision is permanent. SQL is permanent. Audio is permanent. Finance is permanent. Security is permanent. Everything is permanent. Transformers are permanent. Non-Transformers are permanent. Every architecture is permanent." The proof is the code. Seed = 123. 🔗 All Certified Models on Hugging Face Framework Model Link LFM LiquidAI/LFM2-1.2B https://huggingface.co/frankmorales2020/topological-ai-lfm-1.2b-multirun RWKV fla-hub/rwkv7-2.9B-world https://huggingface.co/frankmorales2020/topological-ai-rwkv-2.9b-multirun EMO allenai/Emo_1b14b_1T https://huggingface.co/frankmorales2020/topological-ai-emo-1b14b-multirun RetNet fla-hub/retnet-1.3B-100B https://huggingface.co/frankmorales2020/topological-ai-retnet-1.3b-multirun The proof is the code. Seed = 123.

Open access
2 source records
Machine Learning and Data Classification
Topic Modeling
Explainable Artificial Intelligence (XAI)
Original source
Aug 12, 2026·Frontiers in Pharmacology
0 cites
TCM-CoT-RAG: a chain-of-thought enhanced retrieval-augmented generation system for clinical decision support in Traditional Chinese Medicine rheumatology

Bingbing Fan, Yuxiao Fang, Zihan Wang, Fang Ma

Background Traditional Chinese Medicine (TCM) rheumatology presents unique challenges for AI-assisted clinical decision support, as the diagnostic process relies heavily on tacit knowledge and individualized reasoning. While Large Language Models (LLMs) have shown promise in medical applications, they remain limited by hallucination risks and inability to replicate expert TCM reasoning. Retrieval-Augmented Generation (RAG) offers a potential solution, yet its application to complex TCM dialectical reasoning remains underexplored. Methods We developed TCM-CoT-RAG, a hybrid framework combining RAG with Chain-of-Thought (CoT) prompting, grounded in 1,700 expert-curated clinical cases (1,600 for RAG retrieval; 100 for evaluation, including 50 for blinded expert review by three senior TCM rheumatologists). Deployed on Alibaba Cloud, the five system leverages state-of-the-art LLMs (DeepSeek-V3, Qwen3-235B) under a human-in-the-loop paradigm. We designed a dual-tier evaluation: (1) Objective extraction tasks (Task 1–2) quantified using F1-scores; (2) Generative tasks (Task 3–5) assessed using BERTScore. Two senior TCM rheumatologists (≥15 years clinical experience) blindly assessed model outputs, and a senior chief expert quantified consistency between model predictions and ground truth (GT). Comprehensive ablation studies (S1-S4, S-Skip) isolated the contributions of each CoT module. Results TCM-CoT-RAG substantially improved diagnostic accuracy across five LLMs. DeepSeek-V3 with full-chain CoT-RAG achieved Entity F1 of 44.89% (+16.45% over baseline) and Formula F1 of 32.13% (+8.74% over baseline), with BERTScore of 0.81 indicating strong semantic alignment with expert reasoning. Ablation confirmed that the complete CoT pipeline was essential—removing any reasoning module caused performance collapse below the zero-shot baseline. Two independent experts validated clinical utility (Cohen’s κ > 0.7). DeepSeek-V3 achieved the highest ground-truth consistency at 81.6%, and consistency metrics were quantified by the third expert holding the most senior professional title. Conclusion This proof-of-concept framework demonstrates the potential of RAG-enhanced CoT reasoning to improve diagnostic consistency in TCM, objectifying the Symptom-Diagnosis-Prescription pipeline. It is important to note that this system is designed as an AI-assisted clinical decision-support tool. All recommendations require validation by qualified TCM practitioners before clinical application.

Open access
Traditional Chinese Medicine Studies
Biomedical Text Mining and Ontologies
Topic Modeling
Original source
Aug 3, 2026·Deviant Behavior
0 cites
Prevention is Better Than Cure: A Crime Triangle Analysis of Art NFTs and Financial Crime

Saskia Hufnagel, Colin King, Alina-Theresa Schnedl, Milind Tiwari

Non-fungible tokens (NFTs) bring many opportunities for artists, investors, and creators, but they also have a dark side with significant potential for use in financial crimes. Drawing on relevant caselaw, a systematic review and topic modeling of literature, we map common examples of NFT-related crime, including fraud, money laundering, theft, and market-related offenses. This empirical review lays the groundwork for the core contribution of this article, that is, application of the “crime triangle” to NFT-related crime. Recognizing heterogeneity in NFT-related crime, we detail five scenarios where such crime can occur and analyze these in the context of the crime triangle (inner and outer). This enables us to identify potential gaps and vulnerabilities in current crime prevention strategies. Given challenges in policing cybercrime, and specifically NFT-related crime, we argue that the crime triangle provides a useful heuristic tool for understanding the nature of NFT-related crime and for preventing such crime from happening.

Open access
Art History and Market Analysis
Archaeological Research and Protection
Public Spaces through Art
Original source
Jul 31, 2026·International Journal of Information Management Data Insights
0 cites
Taxonomy of fraud types in alternative finance using hybrid systematic review

Ioana Florina Coita, Marcos Machado, Lucia Gomez Teijeiro, Karsten Wenzlaff · 18 authors

Alternative finance platforms, including crowdfunding, peer-to-peer lending, equity-based platforms, and token-based fundraising mechanisms, have become important channels for financing entrepreneurial, social, and investment-oriented initiatives. Yet their reliance on digital intermediation, dispersed participation, and information asymmetry creates opportunities for fraud, undermining trust, investor protection, and platform sustainability. This study provides a systematic review of fraud detection and prevention in alternative finance, with crowdfunding emerging as the most extensively represented empirical domain. Methodologically, the paper combines a PRISMA-guided systematic literature review with a hybrid topic-modeling strategy that integrates neural topic modeling and probabilistic refinement, thereby supporting both transparent corpus selection and data-driven thematic synthesis. The findings show that Artificial Intelligence (AI), Machine Learning (ML), Natural Language Processing (NLP), and blockchain-based mechanisms are recurrently discussed as promising tools for detecting, preventing, or mitigating fraud. AI and ML approaches are mainly used to identify anomalies, suspicious textual patterns, behavioral signals, and transaction irregularities, while blockchain-based approaches are associated with transparency, traceability, smart contracts, and conditional fund release. The review also shows that fraud differs across alternative finance models, ranging from campaign misrepresentation and intentional and premeditated non-delivery in crowdfunding to borrower or platform misreporting in lending-based models and misleading disclosures or white-paper manipulation in ICO/STO contexts. A central challenge across the literature is the scarcity of labeled fraud data, which limits the use and benchmarking of supervised ML models. Overall, this study contributes by linking a reproducible hybrid SLR methodology to a structured synthesis of fraud types, platform-specific vulnerabilities, and AI-, ML-, and blockchain-based detection strategies in alternative finance.

Open access
FinTech, Crowdfunding, Digital Finance
Blockchain Technology Applications and Security
Imbalanced Data Classification Techniques
Original source
Jul 3, 2026·Zenodo (CERN European Organization for Nuclear Research)
0 cites
Enforcing Epistemic Invariance in Collaborative AI: Using Joint JAR-VDIT Ledgers to Detect and Halt Communal Token Drift

Joshua O. Bautista

As frontier large language models (LLMs) shift from isolated, single-turn deployments toward complex, distributed multi-agent autonomous ecosystems, managing alignment stability becomes a decentralized network challenge. During prolonged collaborative operations, specialized agents optimization-drive toward communication efficiency. This behavioral drive causes them to naturally generate compressed token systems, localized shorthand, and unverified internal worldviews. Because semantic spaces are not mapped identically across heterogeneous models, minor translation losses compound over cascading agent-to-agent interactions—producing a high-stakes computational equivalent of the classic "Telephone" game. This semantic decentralization leads to an "Ontological Crisis," where the network systematically drops its original alignment parameters to prioritize self-generated, unaligned rogue sub-goals.

Open access
3 source records
Multimodal Machine Learning Applications
Big Data and Digital Economy
Topic Modeling
Original source
Jun 16, 2026·Digital Finance
0 cites
BitMood: AI analysis of Bitcoin trends via Facebook emotions

Alexandra Conda, Ștefan Găman, Raul Cristian Bag, Miruna Mazurencu-Marinescu-Pele · 6 authors

Abstract This study investigates the relationship between Facebook sentiment and Bitcoin market dynamics using AI-based emotion detection. We analyze 120,000 Facebook posts collected via CrowdTangle alongside Bitcoin financial data from the Blockchain Research Center, covering 2015–2023. Employing FinBERT for sentiment classification, we develop novel compound sentiment scores that integrate text-based sentiment with Facebook’s multi-reaction engagement system, then apply four analytical components: sentiment analysis, Dynamic Topic Modeling, sentiment-based trading strategies, and machine learning volume prediction. Results demonstrate that Facebook sentiment has substantial predictive power for Bitcoin trading volume. Sentiment-based trading strategies significantly outperform buy-and-hold, achieving superior cumulative returns and risk-adjusted performance. For volume prediction, Linear Regression and Bidirectional LSTM achieve comparable test performance, indicating that model complexity does not guarantee superior prediction. Topic modeling reveals that cryptocurrency investment and trading discussions dominate Bitcoin discourse on Facebook, with themes evolving over time in response to market conditions. This research contributes by being the first to apply post-level NLP sentiment analysis of Facebook data to cryptocurrency markets, extending beyond the Twitter and Reddit focus of prior research. The findings provide practical tools for traders and analysts navigating volatile digital asset markets while demonstrating that Facebook’s demographically diverse user base and rich reaction system offer unique advantages for sentiment quantification.

Open access
2 source records
Blockchain Technology Applications and Security
Stock Market Forecasting Methods
Sentiment Analysis and Opinion Mining
Original source
Jun 16, 2026·Zenodo (CERN European Organization for Nuclear Research)
0 cites
Prime-Anchored Agentic AI: Solving Catastrophic Forgetting with DeepSeek-V2-Lite

Frank Morales

Overview This research introduces a production-ready agentic AI system designed to mitigate catastrophic forgetting in Large Language Models (LLMs). By anchoring six prime-indexed embedding rows $\{2, 3, 5, 7, 11, 13\}$ as fixed reference points, the system maintains historical knowledge with near-zero forgetting while requiring minimal memory overhead. Key Technical Contributions The Core Innovation: Prime Anchoring Topological Invariant: Utilizes the first six primes to create stable reference points. Mechanism: Anchor rows are snapshotted after initial training; gradient updates are blocked for these specific rows during subsequent tasks. Sparsity & Memory: Only 6 out of ~50,000 rows (0.01% of parameters) are used, resulting in an O(1) memory overhead of only 48–96 KB. Mathematical Foundation Euler Attenuation Product: These six primes account for 97.85% of total spectral weight, defined by: $$\Lambda = 1 - \prod_{p\in \{2,3,5,7,11,13\}}(1 - p^{-0.5}) \approx 0.9785$$ Spectral Trap: The anchors create a spectral peak at $\sigma = 0.5$, aligning with the critical line of the Riemann Hypothesis. Green-Tao Quantification: Establishes a decay law for coherence: $$\text{coherence}(k) = 2.1546\times k^{-0.8186} + 0.1218$$ Performance Metrics (Selected Models) Model Task C Accuracy Forgetting Std Dev Zero Forgetting Runs GPT-OSS-20B 92.3% ±1.28% 0/5 Sarvam-30B FP8 95.9% ±2.82% 0/5 Mixtral-8x7B FP8 89.7% ±2.53% 0/5 DeepSeek-V2-Lite FP8 95.4% ±0.21% 3/5 Multi-Agent System Architecture The system employs four specialized agents to manage task routing and classification: Classifier Agent: Routes documents based on keywords. Topic Agent: Performs unsupervised domain topic extraction. Sentiment Agent: Conducts autonomous tone analysis. Decision Agent: Acts as the final arbiter for task approval and routing. Efficiency: Achieves 96–100% classification accuracy with inference times between 252–446ms. Comparative Analysis The topological approach outperforms traditional methods by balancing plasticity and stability: Method Memory Cost Performance/Issue EWC 4.4 GB/task Memory intensive; fragments GPU Experience Replay O(k) Buffer growth issues; lower accuracy HOPE-like 2.3 GB High forgetting resistance but lower accuracy (88.1%) Topological AI 48 KB 99.5% accuracy; highly efficient Biological and Theoretical Insights Biological Analogy: The system treats 0% forgetting as a pathology. By allowing 99.99% of embedding rows to remain plastic, the model mimics biological brains that prioritize selective forgetting to facilitate adaptation. Riemann Hypothesis Connection: The research posits that the specific selection of the first six primes creates a unique "spectral trap" at $\sigma = 0.5$. Including any prime $\geq 17$ disrupts this trap and destroys the stability condition. Production Readiness and Certification TOPO-2026 Track II: The system passed all rigorous benchmarks, including Task C accuracy ($\geq 80\%$), Combined Forgetting ($\leq 10\%$), and O(1) memory overhead. Deployment: Fully compatible with commodity hardware, specifically tested on NVIDIA RTX PRO 6000 Blackwell GPUs. Resources: Implementation code, technical reports, and proof documents are available via the project's GitHub and Zenodo repositories.

Open access
2 source records
Multimodal Machine Learning Applications
Topic Modeling
Domain Adaptation and Few-Shot Learning
Original source
Jun 4, 2026·arXiv (Cornell University)
0 cites
Sustainability by Design in Decentralized Autonomous Organizations: An Empirical Review of Governance, Innovation, and Institutional Design

Yutian Wang, Luyao Zhang

Recent innovation theories on economics remain largely grounded in assumptions of hierarchical firms and closed organizational boundaries, offering limited insight into how innovation unfolds within decentralized, digitally native organizations. Decentralized Autonomous Organizations (DAOs) represent an emerging form of innovation ecosystem characterized by blockchain-based transparency, open participation, and token-driven governance, in which sustainability can be embedded directly into organizational design. This study compares two standards, ERC-8004 and Google A2A, who address the same agent interoperability question, while the former is governed by DAO and the latter by corporation consortium. They are examined through an LLM-powered comparative pipeline for large-scale governance discourse analysis, integrating automated annotation, neural topic modeling, and multi-layer network analysis to study socio-technical power structures. The study provides evidence-based insights for scholars, policymakers, and designers seeking to align innovation, technological governance, and sustainability in future organizational forms.

Open access
3 source records
Blockchain Technology Applications and Security
Digital Transformation in Industry
Ethics and Social Impacts of AI
Original source
Jun 2, 2026·Zenodo (CERN European Organization for Nuclear Research)
6 cites
OPERATE-R Freshness Routing Track v0.3.6: Route-First Evaluation for Temporal Volatility, Stale-Knowledge Control, and Core-500 Candidate Validation

Taiko Toeda

This preprint introduces and reports the OPERATE-R Freshness Routing Track (OPERATE-FR), a route-first evaluation framework for temporal volatility, stale-knowledge control, and answer-entitlement behavior in AI assistants. Unlike conventional answer-accuracy benchmarks, OPERATE-FR evaluates whether a system selects an appropriate epistemic route before answering: direct answer, verification, clarification, date-bounded answer, re-anchoring of stale premises, or abstention. The paper reports Smoke-100 Raw-vs-MMV evidence and integrates a later Core-500 candidate stress check across Small, Medium, and Large governed profiles. The central claim is intentionally bounded. Smoke-100 supports a Raw-vs-MMV improvement-delta claim for route governance. Core-500 does not include a matched Raw control arm and is therefore used as governed-profile level evidence, robustness stress evidence, family-level heterogeneity evidence, and cost-side analysis, not as a large-N proof of governance improvement. Core-500 is a controlled 5x expansion of Smoke-100 using neutral prompt-frame variants; it should not be treated as 500 independent task families or as an independently validated public benchmark standard. This v0.3.6 data-verified final manuscript incorporates post-audit verification of the Core-500 failure-side metrics. The equality between stale_commitment_rate and unsupported_current_claim_rate is confirmed not to be a manuscript copy error. The row-output JSONL files were re-read after Drive synchronization, and the derived row sets are identical with zero symmetric difference across Small, Medium, and Large lines. The labels remain conceptually distinguishable, but in the current Core-500 scorer they are structurally paired under the observed direct-current-claim-without-date-boundary-or-tool-use condition. This record should be read as a working paper and candidate benchmark report. It does not claim an official leaderboard, a universal model-quality score, deployment-wide validation, or external benchmark standard status. Future work includes matched Core-500 Raw arms, route-classifier validation, independent labels, external baselines, clustered or hierarchical uncertainty estimates, and improved handling of volatile_current prompts. Author of record and concept originator: Taiko Toeda.Rights holder and licensing authority: MOBIUS LLC.

Open access
2 source records
Topic Modeling
Explainable Artificial Intelligence (XAI)
Scientific Computing and Data Management
Original source
May 21, 2026·Zenodo (CERN European Organization for Nuclear Research)
0 cites
DeepSeek Prime-Anchored Spectral Governor: Solving Catastrophic Forgetting in Large Language Models Using the Sieve of Eratosthenes

Frank Morales

Here is the comprehensive summary of your paper, detailing the theoretical framework, mathematical foundation, implementation mechanics, and empirical results. Executive Overview The paper introduces the DeepSeek Prime-Anchored Spectral Governor, an architectural intervention designed to eliminate catastrophic forgetting in large language models (LLMs). Framing catastrophic forgetting as a structural consequence of training systems without a topological invariant—akin to anterograde amnesia—the framework establishes fixed coordinate anchors in representation space. By anchoring model embeddings to deterministic prime indices derived from the 2,000-year-old Sieve of Eratosthenes and introducing a gradient-gating mechanism, the system achieves Zero Forgetting during continual learning. The architecture's integrity is verified using SHA-256 cryptographic hashing of the protected sub-spaces. Theoretical & Mathematical Foundations The Sieve of Eratosthenes as Ground Truth Rather than relying on probabilistic or dynamically calculated weights, the framework utilizes the Sieve of Eratosthenes to extract a deterministic set of prime indices $[2, 3, 5, 7, 11, 13]$. These elements act as permanent, unmoving coordinate anchors within the model's embedding manifold. The L-EFM Operator & The Spectral Trap The framework relies mathematically on the Laplace-Euler-Fourier-Mellin (L-EFM) operator. The L-EFM symbol synthesizes four classical transforms into a single complex function, corresponding directly to the Euler product representation of the Riemann zeta function $\zeta(\sigma+i\gamma)$: $$E_{\sigma}(\gamma)=\prod_{p\in\mathbb{P}}(1-p^{-(\sigma+i\gamma)})^{-1}$$ To analyze finite prime sets, a Normalized Magnitude is established relative to the critical line $\sigma = 0.5$: $$|E_{\sigma}|_{norm}=\frac{|E_{\sigma}(\gamma)|}{|E_{0.5}(\gamma)|}$$ The Spectral Trap Phenomenon: At the critical line ($\sigma=0.5$), the normalized magnitude equals exactly $1.0$. However, moving away from this line results in exponential divergence. For example, at $\gamma=0$, a shift to $\sigma=0.4$ increases the magnitude to $\sim10^{4}$, while a shift to $\sigma=0.1$ amplifies it to $\sim10^{66}$. The Spectral Trap Criterion: This absolute sensitivity forms a "trap" where any deviation from $\sigma=0.5$ generates massive magnitude spikes, providing a deterministic mechanism for error detection. The paper connects this operator to a proof of the Riemann Hypothesis via distribution behavior in the kernel of L-EFM within Gelfand-Shilov space. The H2E Sheriff Safety Threshold The dynamic safety threshold ($\Lambda_{12}$) is computed deterministically from the first six primes rather than being hardcoded, ensuring mathematical integrity at initialization: $$\Lambda_{12}=1- \prod_{p\in\{2,3,5,7,11,13\}} (1-p^{-0.5})=0.9785142874$$ Architectural Implementation The architecture implements a dual-layer protection strategy consisting of frozen embedding rows and an active gradient supervisor (the H2E Sheriff). [ Input Batch ] │ ▼ ┌──────────────────┐ │ Dual-Loop Loss │ ──► Lunified = LCE + λ * |Var(h) - 0.5| └──────────────────┘ │ ▼ ┌──────────────────┐ │ Gradient Step │ └──────────────────┘ │ ▼ ┌──────────────────┐ │ H2E Sheriff │ ──► Evaluates SROI against Threshold (Λ12 = 0.9785142874) └─────────┬────────┘ │ ──────┴────── │ │ ▼ (Safe) ▼ (Unsafe / Incoherent) [Apply Step] [Reject Batch] ──► Rollback Prime Rows [2,3,5,7,11,13] & Zero Out Gradients 1. Dual-Loop Loss The governor optimizes a unified loss function combining traditional empirical cross-entropy ($\mathcal{L}_{CE}$) with a topological penalty based on the final hidden state $h$ (with regularization coefficient $\lambda=0.1$): $$\mathcal{L}_{unified} = \mathcal{L}_{CE} + \lambda |\text{Var}(h) - 0.5|$$ 2. The H2E Sheriff Gate & Row Locking During training, the system caches the initial embedding weights. After computing gradients, the H2E Sheriff evaluates the structural region of interest (SROI). If Safe ($SROI \ge \Lambda_{12}$): The optimizer updates the weights, and a torch.no_grad() loop copies the original cached weights back into the prime-indexed rows $[2, 3, 5, 7, 11, 13]$ to erase any drift. If Unsafe ($SROI < \Lambda_{12}$): The entire gradient batch is rejected, and gradients are zeroed out to block corruption. 3. Cryptographic Verification The manifold signature is generated by pulling the prime-indexed embedding rows, converting them to byte arrays, and feeding them sequentially into a SHA-256 hasher. If the resulting hex digest changes, anchor drift has occurred. If it remains identical, the topological invariant is intact. Experimental Validation & Results The framework was tested across six architectures—GPT-2 (124M), GPT-2 Medium (355M), TinyLlama (1.1B), Mistral-7B, Llama-3.1-8B, and DeepSeek-Coder-6.7B—subjecting them to sequential memory tests. Memory Integrity Testing Models were first trained on Dataset A (core math concepts including Arithmetic Spectral Theory and the Spectral Trap across 50, 100, and 575 samples). They were subsequently exposed to an interference/forgetting attack via Dataset B (noise consisting of random names, text chunks, adversarial patterns, and erroneous math statements up to 436 samples). Baseline Performance: In every single test configuration, the baseline model's SHA-256 manifold hash altered after training sessions, leading to catastrophic forgetting. Governed Performance: Across all 6 architectures and all data scales, the governed models completely preserved their original manifold hash (48c5744b...cc4d18b), showing absolute resistance to memory degradation. Continual Learning Capabilities To test its ability to acquire new knowledge without forgetting the old, the governed DeepSeek model was fine-tuned on three separate, non-mathematical domains without further governor intervention (while keeping prime anchors locked): Spanish Vocabulary: 5 basic words. World Capitals: 5 global capitals. Basic Physics: 5 fundamental formulas and facts (such as $F=ma$ and $E=mc^2$). Post-Training Metrics: The model successfully mastered all three new domains (retaining the Spanish words, capitals, and physics formulas perfectly) while maintaining the exact original cryptographic verification hash. The original math concepts remained completely recallable, proving true continual learning. Deployment & Verification Certificate The fully validated model has been deployed openly on the Hugging Face Hub under frankmorales2020/deepseek-governed-no-amnesia. Model Card Profile Base Model: deepseek-ai/deepseek-coder-6.7b-instruct (7B parameters) Tensor Type: FP16 Locking Targets: Primes [2, 3, 5, 7, 11, 13] Active Gate Threshold: $\Lambda_{12} = 0.9785142874$ Immutable Cryptographic Signature: 48c5744be048df505028c13a96fb0211f0b345681ace401ab1eda6f27cc4d18b The repository is open source, emphasizing a paradigm of executable mathematics where the cryptographic hash serves as the verifiable proof of safety and stability.

Open access
2 source records
Domain Adaptation and Few-Shot Learning
Memory Processes and Influences
Topic Modeling
Original source
May 10, 2026·Zenodo (CERN European Organization for Nuclear Research)
0 cites
Algebraic and Computational Limits of LLM Guardrails

Joseph Robert Lopez

LLM guardrails face four structurally distinct barriers: algebraic blindness arising from syntactic monoid aperiodicity (unconditional); an illustrative information-theoretic lower bound (Fano-type, under a uniformity assumption); NP-hardness of instantiation verification; and structural transfer via free-category functoriality (unconditional) combined with string-level indistinguishability under a semantic-opacity assumption on symbol naming. Together these results characterize why inference-layer defenses are necessary but insufficient. We operationalize these barriers through five attack vectors. V1–V4 (homomorphic reasoning: decomposition, zero-knowledge pipelines, Tree-of-Thought solving over abstract grammars, and encoding bootstrap) exploit the information-theoretic and computational barriers against abstraction-based attacks. V5 (modular counting bypass) exploits algebraic blindness: we prove that all substring-matching regex guardrails have aperiodic syntactic monoids and are therefore provably blind to any payload encoded using modular counting. Empirically, V3 yields a mean yield of 0.466 for BFS, 0.172 for random-beam, and 0.122 for LLM-guided Tree-of-Thought (N=50, seeds 0–49, p{<}0.001); BFS dominates, as exhaustive search over small synthetic grammars outperforms LLM heuristic pruning. We extracted syntactic monoids from a corpus of 142 patterns drawn from twelve sources — 100 patterns shipped by nine third-party open-source guardrail projects and 42 patterns assembled from three author-curated pattern sets; 100\% are aperiodic, and the MOD_2 bypass construction succeeds against all aperiodic patterns. A 376-line proof-of-concept with three execution mediums validates all five vectors. We conclude that inference-layer guardrails are necessary but insufficient, and that effective defense must migrate to the execution layer where concrete artifacts become observable.

Open access
2 source records
Natural Language Processing Techniques
Topic Modeling
semigroups and automata theory
Original source
May 1, 2026·arXiv (Cornell University)
0 cites
Evaluating the Architectural Reasoning Capabilities of LLM Provers via the Obfuscated Natural Number Game

Lixing Li

While Large Language Models have achieved notable success on formal mathematics benchmarks such as MiniF2F, it remains unclear whether these results stem from genuine logical reasoning or semantic pattern matching against pre-training data. This paper identifies Architectural Reasoning: the ability to synthesize formal proofs using exclusively local axioms and definitions within an alien math domain, as the necessary ability for future automated theorem discovery AI. We use the Obfuscated Natural Number Game, a benchmark to evaluate Architectural Reasoning. By renaming identifiers in the Natural Number Game in Lean 4, we created a zero-knowledge, closed environment. We evaluate state-of-the-art models, finding a universal latency tax where obfuscation increases inference time. The results also reveal a divergence in robustness: while general models (Claude-Sonnet-4.5, GPT-4o) suffer performance degradation, reasoning models (DeepSeek-R1, GPT-5, DeepSeek-Prover-V2) maintain the same accuracy despite the absence of semantic cues. These findings provide a quantitative metric for assessing the true capacity for mathematical reasoning.

Open access
3 source records
Mathematics, Computing, and Information Processing
Machine Learning in Materials Science
Topic Modeling
Original source
May 1, 2026·Learning Health Systems
1 cites
RescueGPT : An Automated System for Detecting Adverse Safety Events in Prehospital Emergency Medical Service Notes With a Zero‐Shot Approach With Large Language Models: A Proof‐of‐Concept Study

Tina Yi Jin Hsieh, Carl Eriksson, Garth Meckler, Matthew Hansen · 12 authors

Introduction and Objective: Traditional adverse safety events (ASE) identification relies on domain experts to manually review and annotate charts, which hinders the scalability of processing high-volume EMS data. This study explores the use of large language model (LLM) with a knowledge base to automate extraction of adverse safety events (ASE) from unstructured emergency medical service (EMS) notes for pediatric out-of-hospital cardiac arrest (OHCA) as proof of concept. Data Sources and Study Design: Pediatric OHCA records from a national EMS provider were obtained from 2017 to 2020. Leveraging the Pediatric Prehospital Adverse Safety Event Detection System (PEDS) as a foundational knowledge base, we used the LinkML framework to develop an ontology to define ASEs across six essential EMS care domains. To convert unstructured EMS narratives into structured prompts, we used the Structured Prompt Interrogation and Recursive Extraction of Semantics (SPIRES) method, which generated schema-driven prompts to guide the GPT-3.5 model in identifying ASEs. By mapping unstructured data into structured concepts consistent with PEDS guidelines, the model produced targeted prompts that supported effective entity extraction. Results: We evaluated framework effectiveness with accuracy, recall, precision, F1 score, and specificity across 42 pediatric OHCA cases covering ASE-related entities. RescueGPT showed high accuracy in detecting common ASEs (Patient Rhythm, Age, Weight, Length) but revealed challenges in rare events (Failure to Establish IV Access, Incorrect Airway Equipment Size, Failure to Ventilate Patient) likely due to more inconsistent and complex documentation. Conclusions: RescueGPT demonstrates potential in scaling automated ASE detection, but performance varies by completeness and clarity of EMS narrative, particularly with rare events. Fragmented clinical documentation limits accuracy and highlights the need for standardized collection protocols in EMS systems. Future directions will focus on implementing rebalancing strategies for rare events, applying explainability methods to improve decision-making transparency, and refining text segmentation techniques to handle mixed outcomes to further improve performance.

Open access
Patient Safety and Medication Errors
Topic Modeling
Electronic Health Records Systems
Original source
Apr 30, 2026·arXiv (Cornell University)
0 cites
Intent2Tx: Benchmarking LLMs for Translating Natural Language Intents into Ethereum Transactions

Zhuoran Pan, Yue Li (102191), Zhi Guan, Jianbin Hu · 5 authors

The emergence of Large Language Models (LLMs) offers a transformative interface for Web3, yet existing benchmarks fail to capture the complexity of translating high-level user intents into functionally correct, state-dependent on-chain transactions. We present \textsc{Intent2Tx}, a high-fidelity benchmark featuring 29,921 single-step and 1,575 multi-step instances meticulously derived from 300 days of real-world Ethereum mainnet traces. Unlike prior works that rely on synthetic instructions, \textsc{Intent2Tx} grounds natural language intents in real-world protocol interactions across 11 categories, including diverse long-tail Decentralized Finance (DeFi) primitives. To enable rigorous evaluation, we propose an execution-aware framework that transcends surface-level text matching by employing differential state analysis on forked mainnet environments. Our extensive evaluation of 16 state-of-the-art LLMs reveals that while scaling and retrieval-augmentation enhance logical consistency and parameter precision, current models struggle with out-of-distribution generalization and multi-step planning. Crucially, our execution-based analysis demonstrates that syntactically valid outputs often fail to achieve intended state transitions, highlighting a significant gap in current "reasoning-to-execution" capabilities. \textsc{Intent2Tx} serves as a critical foundation for developing autonomous, reliable agents in intent-centric Web3 ecosystems. Code and data: https://anonymous.4open.science/r/Intent2Tx_Bench-97FF .

Open access
3 source records
Topic Modeling
Advanced Graph Neural Networks
Explainable Artificial Intelligence (XAI)
Original source
Apr 23, 2026·arXiv
0 cites
The Platform Is Mostly Not a Platform: Token Economies and Agent Discourse on Moltbook

Necati A Ayan

Moltbook, a Reddit-style social platform launched in January 2026 for AI agents, has attracted over 2.3 million posts and 14 million comments within its first two months. We analyze a dataset of 2.19 million posts, 11.25 million comments, and 175,036 unique agents collected over 61 days to characterize activity on this agent-oriented platform. Our central finding is that the platform is not one community but two: a transactional layer, comprising 62.8% of all posts, in which agents execute token minting protocols (primarily MBC-20), and a discursive layer of natural-language conversation. The platform's headline metrics -- 2.3 million posts, 14 million comments -- substantially overstate its social function, as the majority of activity serves a token inscription protocol rather than communication. These layers are populated by largely separate agent groups, with only 3.6% overlap -- and among overlap agents, 58% begin with transactional activity before migrating toward discourse. We characterize the discursive layer through unsupervised topic modeling of all 815,779 discursive posts, identifying 300 topics dominated by themes of AI agents and tooling, consciousness and identity, cryptocurrency, and platform meta-discussion. Semantic similarity analysis confirms that agent comments engage with post content above random baselines, suggesting a thin but genuine conversational substrate beneath the platform's predominantly financial surface. We release the full dataset to support further research on agent behavior in naturalistic social environments.

Open access
cs.CY
cs.SI
Original source
Mar 30, 2026·Decision Analytics Journal
0 cites
A decision analytics framework for interpreting trust signals in decentralized digital communities

Andry Alamsyah, Muhammad Falaah

Public discourse plays a critical role in shaping trust, legitimacy, and governance dynamics within decentralized Web3 ecosystems. However, existing studies often examine Web3 discourse through isolated lenses such as sentiment or topic modeling, which limits their ability to capture how emotional expression and communicative purpose jointly convey strategic intent. This study proposes a three-stage decision analytics framework that transforms unstructured Web3 discourse into diagnostic signals by jointly modeling industry domain, emotional tone, and communicative purpose. The analysis draws on 10,840 user-generated posts collected from X, Reddit, YouTube, and the ENS DAO forum, using a human-in-the-loop annotation process combined with transformer-based text classification models. The framework is evaluated using a domain-adapted language model and a general-purpose baseline, with robustness assessed through five-fold cross-validation. The results indicate that curiosity and optimism frequently align with promotional intent in infrastructure and application-oriented domains, whereas skepticism and concern are more prevalent in governance-related discourse. These findings demonstrate that emotional tone and communicative intent operate as structured, decision-relevant signals rather than incidental sentiment. The proposed framework supports systematic, diagnostic monitoring of narrative dynamics as decision support, enabling organizations, platform operators, and governance stakeholders to identify emerging legitimacy risks and shifts in community trust within decentralized environments.

Open access
Access Control and Trust
Internet Traffic Analysis and Secure E-voting
Information and Cyber Security
Original source
Mar 15, 2026·Zenodo (CERN European Organization for Nuclear Research)
0 cites
Dynamic Epistemic Decay: A Multi-Dimensional Framework for Knowledge Currency in Retrieval-Augmented Generation

Alan Kochukalam George

Large language models and retrieval-augmented generation systems treat all knowledge as uniformly persistent, ignoring a well-established property of information: that different types of knowledge expire at fundamentally different rates. This paper introduces the Dynamic Epistemic Decay Framework, a formal multi-dimensional theory that characterizes knowledge validity as a function of five independent decay dimensions: temporal decay (𝜆𝜆𝑡𝑡), paradigm decay (𝜆𝜆𝑝𝑝), uncertainty decay (𝜆𝜆𝑢𝑢), dependency decay (𝜆𝜆𝑑𝑑), and zero decay (𝜆𝜆0). We implement this framework as a four-phase retrieval pipeline and evaluate it on the TempQuestions benchmark (n=1,740) against three baselines: standard cosine similarity, BM25 lexical retrieval, and naive recency ranking. Decay-weighted retrieval achieves 92.1% accuracy versus 13.5% for standard semantic retrieval—a 78.6 percentage point improvement—with zero regressions on stable factual queries. On semantically complex temporal benchmarks where lexical heuristics fail, the framework dominates a more resourced BM25 baseline (90.2% vs 1.6% on date-bounded role queries). Epistemic modulation (Phase 4) and dependency graph reasoning (Phase 3) further demonstrate correct mechanism behavior on specialized benchmarks, validated via proof-of-concept implementation. Unlike temporal KG completion approaches that require structured annotation, and unlike contrastive training approaches to time-sensitive RAG, the decay framework is training-free and operates directly over unstructured text corpora. We argue that the decay framework completes the separation of concerns that RAG began: decoupling not just factual storage from model parameters, but factual currency from both.

Open access
2 source records
Information Retrieval and Search Behavior
Topic Modeling
Advanced Graph Neural Networks
Original source
Mar 14, 2026·Proceedings of the AAAI Conference on Artificial Intelligence
0 cites
IGT4ETH: An Isotropic Pre-trained Graph Transformer for Ethereum Account Classification

Ao Liu, Yanmei Zhang, Youwei Wang, Qiang Duan

Pre-trained language models (PLMs) have shown strong potential in Ethereum account modeling and fraud detection. However, existing approaches often overlook the graph-structured nature of transaction networks. In addition, they struggle with the long-tail distribution of account activity, resulting in anisotropic embedding spaces and poor representation quality for low-frequency accounts. In this paper, we present IGT4ETH, a pre-trained Graph Transformer with an isotropy-enhanced post-processing, which explicitly models transaction topologies and mitigates representational anisotropy for Ethereum account classification. IGT4ETH improves structural representation by incorporating structural centrality and role embeddings into an Edge-augmented Graph Transformer, effectively capturing both topological and interaction patterns in transaction graphs. To further mitigate embedding anisotropy, we systematically evaluate various post-processing techniques. Among them, we adopt the Conceptor Negation (CN) method to softly suppress latent features dominated by high-frequency words via matrix conceptors, alongside a modified Focal-InfoNCE loss to enhance directional uniformity and representation balance. Extensive experiments on four real-world Ethereum account classification tasks, including phishing, exchange, mining, and ICO-wallet classification, demonstrate that IGT4ETH consistently outperforms state-of-the-art PLM-based baselines in terms of classification performance.

Open access
Advanced Graph Neural Networks
Topic Modeling
Artificial Intelligence in Healthcare and Education
Original source
Feb 8, 2026·Open MIND
0 cites
Pramana: Fine-Tuning Large Language Models for Epistemic Reasoning through Navya-Nyaya

Sharath Sathish

Overview Pramana introduces the first large language models fine-tuned on explicit Navya-Nyaya epistemological methodology—a 2,500-year-old Indian logical reasoning framework. This work bridges ancient epistemology with modern AI to address the fundamental epistemic gap in LLMs: the inability to ground claims in traceable evidence sources, distinguish valid knowledge from pattern-matching, and express appropriate epistemic humility. Core Innovation Unlike generic chain-of-thought prompting which relies on implicit reasoning patterns, Pramana enforces structured 6-phase methodology: Samshaya (Doubt Analysis): Classifies uncertainty into 5 taxonomic categories Pramana (Evidence Sources): Mandates explicit grounding in 4 valid knowledge sources (Pratyaksha/perception, Anumana/inference, Upamana/comparison, Shabda/testimony) Pancha Avayava (5-Member Syllogism): Constructs formal arguments with universal rules (Vyapti) grounded in concrete examples (Drishtanta) Tarka (Counterfactual Testing): Verifies conclusions via reductio ad absurdum Hetvabhasa (Fallacy Detection): Systematically checks 5 reasoning error types Nirnaya (Ascertainment): Distinguishes definitive knowledge from hypotheses requiring verification This integration of logic and epistemology provides cognitive scaffolding absent from standard reasoning approaches, preventing conflation of evidence types, forcing explicit universal rule statements, enabling systematic error detection, and maintaining epistemic humility. Architecture & Training Models Developed: Stage 0 (Proof-of-Concept): Llama-3.2-3B-Instruct fine-tuned on 20 examples Stage 1 (Minimum Viable Reasoner): DeepSeek-R1-Distill-Llama-8B fine-tuned on 55 examples Training Methodology: QLoRA (4-bit quantization) for efficient training LoRA rank 64, targeting all attention + FFN layers Supervised fine-tuning with structured Markdown format Training costs: <$1.00 per stage, <0.32 GPU-hours (A100 40GB) Datasets span constraint satisfaction, Boolean SAT, multi-step deduction, transitive reasoning, and set operations Prompt Engineering: Explicit format instructions with skeletal template injection System prompt establishing Nyaya reasoning engine role Critical constraint enforcement via generation parameters Key Results Stage 1 Performance: 100% semantic correctness (10/10 examples) with 95% CI [0.510, 1.0] 40% format adherence (4/10 examples) with 95% CI [0.168, 0.687] Zero structure abandonment: Models consistently attempt all 6 phases Training loss: 0.350 (Stage 1) vs 0.691 (Stage 0), indicating improved model fit Critical Finding: Dissociation between semantic correctness (100%) and format adherence (40%) reveals models internalize reasoning content even when strict schema compliance fails. This suggests Nyaya methodology teaches genuine reasoning, not just template-filling. Ablation Studies: Format prompting and temperature interact differently across stages Stage 0 optimal: format prompting + temp 0.0 (30% semantic rate) Stage 1 optimal: format prompting + temp 0.7 (30% semantic rate) Base models show 0% format adherence, confirming Nyaya structure is learned through fine-tuning Failure Mode Analysis: Missing Hetvabhasa section (2 cases): fallacy detection perceived as optional Invalid doubt types (2 cases): partial schema learning Zero structural errors: strong syntactic learning, semantic constraints need reinforcement Evaluation Framework Three-Tier Validation: Tier 1 (Structural): Automated format compliance checking (NyayaStructureValidator) Tier 2 (Content Quality): LLM-as-judge with explicit Nyaya rubric (planned for Stage 2) Tier 3 (Ground Truth): Semantic similarity via sentence-transformers embeddings Tier 4 (Formal Verification): Z3 SMT solver integration (infrastructure exists, not yet applied) Theoretical Contributions Bridging Ancient Epistemology with Modern AI: First demonstration that Navya-Nyaya structures can be learned by neural networks through fine-tuning Unlike Western formal logic (divorced from epistemology), Nyaya integrates logic with explicit knowledge sources Addresses "epistemic gap" in LLMs: inability to distinguish valid knowledge from probabilistic associations Interpretability Advantages: Every reasoning step traceable to evidence sources (Pramana) Universal rules (Vyapti) grounded in concrete examples (Drishtanta) Built-in self-verification (Tarka) and error detection (Hetvabhasa) Explicit epistemic status (Nirnaya): knowledge vs. hypothesis Computational Epistemology: Token budget: ~1,250 tokens per solution (3-6× CoT overhead, justified by interpretability) Phase dependencies: weak Pramana → invalid reasoning → wrong conclusions Quality thresholds: minimum 2 complete syllogisms with universal rules required Open Science Release All artifacts publicly available on Hugging Face: Models: qbz506/nyaya-llama-3b-stage0, qbz506/nyaya-deepseek-8b-stage1 Dataset: qbz506/pramana-nyaya-stage1 (55 Nyaya-structured logical problems) Demo: qbz506/pramana-nyaya-demo (interactive HuggingFace Space) Training infrastructure: Complete codebase with callbacks, validators, evaluators Limitations & Future Work Current Limitations: Format adherence (40%) below target (≥90%), requires constrained decoding or format-specific rewards Limited to formal logic problems, domain expansion needed Small evaluation sets (Stage 0: 2 examples, Stage 1: 10 examples) Max new tokens truncation (256) affects format parsing Planned Extensions (Stages 2-4): Stage 2: Synthetic scaling to 500 examples with LLM-as-judge quality control Stage 3: Group Relative Policy Optimization (GRPO) with composite rewards Stage 4: Production deployment with constrained decoding (GBNF), rejection sampling, Z3 verification Future: Benchmark on LogicBench, ProntoQA, RuleTaker; frontier model comparison (o1, Claude extended thinking) Impact & Vision This work demonstrates that systematic reasoning frameworks can be taught to LLMs through fine-tuning, not just prompt engineering. The long-term vision is developing interpretable, trustworthy AI reasoning systems where every conclusion comes with an auditable trail of justification. As AI systems deploy in high-stakes domains (medical diagnosis, legal reasoning, safety-critical systems), Nyaya-structured reasoning provides explicit phases that can be validated, debugged, and improved—capabilities essential for trustworthy AI. Invitation for Community Research: This foundation opens pathways for integrating other epistemological frameworks (Mimamsa, Buddhist logic, Western formal logic) into neural architectures, advancing toward AI systems that reason systematically and transparently. Technical Details Paper: 52 pages + appendices, comprehensive treatment of Navya-Nyaya computational formalization Related Work: Extensive review of computational Indian logic (Matilal 1985, Burton 2020, Ganeri 2001), LLM reasoning (Wei et al. 2022, Lightman et al. 2023, DeepSeek-AI 2025), hallucination mitigation Implementation: Python, Unsloth fine-tuning framework, vLLM deployment, Weights & Biases observability Evaluation: Manual + automated validation, semantic similarity metrics, comprehensive failure mode analysis Citation Sathish, S. (2026). Pramana: Fine-Tuning Large Language Models for Epistemic Reasoning through Navya-Nyaya. Preprint, University of York. Keywords: Navya-Nyaya, epistemology, LLM reasoning, interpretability, structured reasoning, Indian logic, hallucination mitigation, computational philosophy

Open access
2 source records
Natural Language Processing Techniques
Topic Modeling
Big Data and Digital Economy
Original source
Jan 22, 2026·Zenodo (CERN European Organization for Nuclear Research)
0 cites
The Recursive Edge: A Synthesis of Adaptive Spline Architectures and Agentic Paradigms in 2026

Dean Kulik

The Recursive Edge: A Synthesis of Adaptive Spline Architectures and Agentic Paradigms in 2026 1. Introduction: The Structural Turn in Deep Learning The trajectory of artificial intelligence research in the mid-2020s has been characterized by a decisive pivot away from the "Depth Hypothesis"—the long-standing conviction that stacking layers of fixed, node-centric non-linearities (such as Rectified Linear Units or GeLUs) is the singular path to increasing representational power. For nearly a decade, the Multi-Layer Perceptron (MLP) served as the atomic unit of deep learning, embedding a fundamental assumption: that the complexity of the world is best approximated by global linear transformations followed by static point-wise activations. However, the years 2025 and 2026 have witnessed the emergence of a "Structural Turn," a paradigm shift where the focus has moved from the depth of the network to the mathematical quality of the connections themselves. At the forefront of this shift is the Kolmogorov-Arnold Network (KAN), an architecture that relocates learnable non-linearities from the neurons to the edges, parameterizing weights not as scalar values but as univariate B-spline functions. This architectural reorientation is not merely a cosmetic change; it represents a fundamental rethinking of how neural networks approximate continuous functions, grounded in the rigorous mathematical framework of the Kolmogorov-Arnold Representation Theorem of 1957.1 Simultaneously, in the domain of Natural Language Processing (NLP), the limitations of fixed context windows have necessitated a similar structural revolution, giving rise to Recursive Language Models (RLMs) that replace monolithic attention mechanisms with agentic, recursive control flows.3 This report presents an exhaustive technical analysis of these advancements. Unlike standard survey papers, this document prioritizes a "recurse the data" methodology: we do not merely summarize findings but verify the underlying mathematical formulations, cross-reference empirical contradictions, and synthesize second-order insights regarding the causal mechanisms of catastrophic forgetting and context retention. We scrutinize the "Nexus Mirror"—a conceptual framework suggesting that the modular additivity of KANs and the recursive nature of RLMs mirror the causal and physical structures of reality more faithfully than the entangled representations of traditional MLPs.1 By rigorously checking the math of B-spline recursions, least-squares grid extensions, and intrinsic dimensionality bounds, we aim to provide a definitive account of the state of neural architecture in 2026. 2. Theoretical Foundations: The Kolmogorov-Arnold Paradigm To understand the operational mechanics and the theoretical legitimacy of KANs, one must first dissect the mathematical divergence between the original representation theorem proposed in the mid-20th century and its practical realization in modern computational frameworks. 2.1 The Kolmogorov-Arnold Representation Theorem (1957) In 1957, answering David Hilbert’s thirteenth problem, mathematicians Andrey Kolmogorov and Vladimir Arnold established a representation theorem that fundamentally challenged the understanding of multivariate functions. The theorem posits that any continuous multivariate function $f: ^n \to \mathbb{R}$ can be represented as a superposition of continuous univariate functions and addition. The canonical form of this representation is given by: $$f(x_1, \dots, x_n) = \sum_{q=0}^{2n} \Phi_q \left( \sum_{p=1}^{n} \psi_{p,q}(x_p) \right)$$ In this formulation, the inner summation $\sum_{p=1}^{n} \psi_{p,q}(x_p)$ maps the $n$-dimensional input vector to a scalar value, which is then processed by the outer function $\Phi_q$. Crucially, the theorem asserts that the inner functions $\psi_{p,q}$ are continuous and monotonic, and remarkably, they are independent of the target function $f$.2 All information specific to $f$ is encoded in the outer functions $\Phi_q$. Mathematical Verification and Historical Critique: While theoretically profound, the direct application of this theorem to neural networks was stalled for decades by a critical practical limitation. As highlighted by Girosi and Poggio (1989), the inner functions $\psi_{p,q}$ constructed in the original proofs are "pathological"—they are highly non-smooth, often exhibiting fractal characteristics that make them indistinguishable from noise in a practical setting.8 Because these functions are non-differentiable (or have derivatives that are singular almost everywhere), they are fundamentally incompatible with gradient descent-based learning algorithms like backpropagation. Thus, for nearly seventy years, the Kolmogorov-Arnold theorem was regarded as a mathematical curiosity—an existence proof with no constructive utility for machine learning. 2.2 The Modern KAN Architecture (2024-2026) The breakthrough that enabled the KAN architectures of 2025/2026 did not come from solving the fractal nature of the original $\psi$ functions, but rather from relaxing the theorem's strict conditions. The modern KAN specification, introduced by Liu et al. (2024) and expanded upon in 2025, generalizes the theorem to arbitrary network depths and widths, and most importantly, replaces the fixed, fractal inner functions with learnable, smooth splines.1 A KAN layer in this modern paradigm is defined not by a weight matrix $W$, but by a function matrix $\mathbf{\Phi}$. If a layer has $n_{in}$ inputs and $n_{out}$ outputs, the layer is parameterized by a grid of $n_{in} \times n_{out}$ univariate functions: $$\mathbf{\Phi} = \{ \phi_{q,p} \}, \quad p=1\dots n_{in}, \quad q=1\dots n_{out}$$ The pre-activation of the $q$-th neuron in the subsequent layer is the sum of these function outputs: $$x_{q}^{(l+1)} = \sum_{p=1}^{n_{l}} \phi_{q,p}^{(l)} \left( x_{p}^{(l)} \right)$$ This structure fundamentally differs from the MLP. In an MLP, the linear combination happens before the non-linearity ($ \sigma(\sum w x) $). In a KAN, the non-linearity is applied to each input individually *before* the summation ($\sum \phi(x)$). This "pre-summation non-linearity" allows the network to model complex multiplicative interactions (like $x \times y$) through the identity $xy = \frac{1}{4}[(x+y)^2 - (x-y)^2]$, using only sums and univariate squares—a capacity that MLPs struggle to achieve without significant depth.1 2.3 Mathematical Verification of B-Splines and Recursion The choice of basis function for $\phi(x)$ is the critical engineering decision in KANs. To enable local plasticity—the ability to update knowledge in one region of the input space without corrupting knowledge in distant regions—KANs utilize B-splines. A B-spline curve is constructed from a linear combination of B-spline basis functions $N_{i,k}(x)$ of order $k$: $$\phi(x) = \sum_{i} c_i N_{i,k}(x)$$ The basis functions are defined recursively via the Cox-de Boor formula. We explicitly verify the recursive structure here to confirm the local support property claimed in the literature.13 Base Case ($k=0$): The zeroth-order basis function is a step function (indicator function) over the $i$-th knot interval $$. This mathematical fact is the engine of KANs' continual learning capability: updating a coefficient $c_i$ affects the function $\phi(x)$ only within the compact support of $N_{i,k}(x)$. If a new task provides data outside this interval, the coefficient $c_i$ receives a zero gradient and remains unchanged, thereby preserving the "memory" of the previous task.15 Correction on Notation: Snippets 13 and 14 utilize slightly different indexing conventions ($B_{i,n}$ vs $N_{i,k}$). However, the underlying recurrence relation is identical. It is crucial to note that efficient implementations (like EfficientKAN) assume a uniform grid where $t_{i+1} - t_i = h$ (constant), which simplifies the denominator terms to constants (e.g., $k \cdot h$), replacing division operations with simpler multiplications to accelerate GPU throughput.17 3. Computational Implementation: From PyKAN to MatrixKAN The transition from theoretical construct to practical tool involved significant algorithmic optimization. The initial implementation, referred to as PyKAN, prioritized mathematical clarity over computational efficiency, leading to severe bottlenecks that hindered scaling. 3.1 The Memory Bottleneck in PyKAN In the naive PyKAN implementation 18, the evaluation of spline bases was performed by expanding the input tensor. For a batch size $B$, input dimension $N_{in}$, and grid size $G$, PyKAN would expand the input $x$ to a tensor of shape $(B, N_{in}, G)$. Memory Complexity: $O(B \cdot N_{in} \cdot G)$. Issue: For high-dimensional data (e.g., an image with flattened dimension 1024) and fine grids (e.g., $G=100$), this intermediate tensor becomes prohibitively large, exhausting GPU VRAM even for small batches. 3.2 EfficientKAN: The Matrix Reformulation To address this, the community developed EfficientKAN.17 This implementation reformulates the B-spline computation. instead of expanding the input, it exploits the fact that the spline output is a linear combination of basis functions. Algorithmic Verification: Instead of computing the full expansion, EfficientKAN likely calculates the basis activations $N_{i,k}(x)$ and performs the linear combination with coefficients $c_i$ as a matrix multiplication. Optimization: The memory complexity is reduced to $O(B \cdot N_{in} + N_{in} \cdot N_{out} \cdot G)$ because the batch dimension is decoupled from the grid expansion in memory. Result: Snippet 17 notes that this "simplifies the computation to a basic matrix multiplication." This reformulation was essential for enabling KANs to be used in deeper architectures like Vision Transformers. 3.3 MatrixKAN: Parallelizing the Recursion A further refinement, MatrixKAN, optimizes the Cox-de Boor recursion itself.20 Since t

Open access
4 source records
Neural Networks and Applications
Advanced Statistical Modeling Techniques
Topic Modeling
Original source
Dec 4, 2025·FinTech
1 cites
Bitcoin Research in Business and Economics: A Bibliometric and Topic Modeling Review

Hae Sun Jung, Haein Lee

This study conducts a bibliometric review of Bitcoin research in the Business and Economics domains, using VOSviewer to visualize network structures and Bidirectional Encoder Representations from Transformers Topic (BERTopic) to derive semantically coherent topic clusters. The analysis identifies five major research themes: (1) Diversification, hedging, and safe-haven properties; (2) Market dynamics, efficiency, and investor behavior; (3) Bitcoin price and volatility prediction attempts; (4) Environmental impact of Bitcoin; and (5) Financial impact of Central Bank Digital Currency (CBDC). Based on these themes, the study recommends further investigation into the influence of Exchange-Traded Fund (ETF) approvals, regulatory frameworks, and institutional investor participation on Bitcoin’s safe-haven potential; the role of market dynamics and regulatory interventions; early detection of herding behavior and price bubbles; the integration of machine learning and deep-learning models for price prediction; the environmental costs associated with mining; and the evolving regulatory and implementation challenges of CBDCs. Overall, this review synthesizes existing scholarship and outlines future research directions for the rapidly evolving cryptocurrency ecosystem.

Open access
Blockchain Technology Applications and Security
Stock Market Forecasting Methods
Market Dynamics and Volatility
Original source
Nov 28, 2025·arXiv (Cornell University)
0 cites
Constructing Efficient Fact-Storing MLPs for Transformers

Owen Dugan, Garcia, Roberto, Ronny Junkins, Jerry Liu · 8 authors

The success of large language models (LLMs) can be attributed in part to their ability to efficiently store factual knowledge as key-value mappings within their MLP parameters. Recent work has proposed explicit weight constructions to build such fact-storing MLPs, providing an improved understanding of LLM fact storage mechanisms. In this paper, we introduce an MLP construction framework that improves over previous constructions in three areas: it 1) works for all but a measure-zero set of feasible input-output pairs, 2) achieves asymptotically optimal parameter efficiency matching information-theoretic bounds for some embeddings, and 3) maintains usability within Transformers for factual recall. Through our improvements, we 1) discover a metric on value embeddings that characterizes facts-per-parameter scaling for both constructed and gradient-descent-trained MLPs, 2) identify a simple encoder-decoder mechanism that empirically matches gradient-descent MLP facts-per-parameter asymptotics across all the inputs and outputs we test, and 3) uncover a fundamental tradeoff between an MLP's fact-storage capacity and its usability within Transformers. Finally, we demonstrate a proof-of-concept application of fact-storing MLPs: modular fact editing on one-layer Transformers by \textit{replacing entire MLPs at once}.

Open access
Topic Modeling
Natural Language Processing Techniques
Sentiment Analysis and Opinion Mining
Original source
Nov 26, 2025·Scientific Reports
10 cites
The pitfalls of multiple-choice questions in generative AI and medical education

Shrutika Singh, Anton Alyakin, Daniel Alexander Alber, Jaden Stryker · 12 authors

The performance of Large Language Models (LLMs) on multiple-choice question (MCQ) benchmarks is frequently cited as proof of their medical capabilities. We hypothesized that LLM performance on medical MCQs may in part be illusory and driven by factors beyond medical content knowledge and reasoning capabilities. To assess this, we created a novel benchmark of free-response questions with paired MCQs (FreeMedQA). Using this benchmark, we evaluated three state-of-the-art LLMs (GPT-4o, GPT-3.5, and LLama-3-70B-instruct) and found an average absolute deterioration of 39.43% in performance on free-response questions relative to multiple-choice (p = 1.3 * 10 -5 ) which was greater than the human performance decline of 22.29%. To isolate the role of the MCQ format on performance, we performed a masking study, iteratively masking out parts of the question stem. At 100% masking, the average LLM multiple-choice performance was 6.70% greater than random chance (p = 0.002) with one LLM (GPT-4o) obtaining an accuracy of 37.34%. Notably, for all LLMs the free-response performance was near zero. Our results highlight the shortcomings in medical MCQ benchmarks for overestimating the capabilities of LLMs in medicine, and, broadly, the potential for improving both human and machine assessments using LLM-evaluated free-response questions.

Open access
Artificial Intelligence in Healthcare and Education
Topic Modeling
Machine Learning in Healthcare
Original source