Abstract Environmental transition is increasingly governed through multilevel systems in which authority is shared across supranational, national, and regional governments. Existing research on multilevel climate governance has focused on coordination, implementation, and compliance, largely treating environmental objectives as politically consistent among territorial levels once adopted. This paper argues that multilevel governance also reshapes the political content of environmental transition itself, a process we call political reinterpretation: common climate objectives are selectively reprioritized and reframed as they enter territorially distinct political arenas. We test this argument using text analysis of parliamentary discourse, applying structural topic modeling to 10,564 speeches delivered across Spain’s seventeen Autonomous Communities (ACs) between 2019 and 2024 to uncover five substantive dimensions of environmental-transition discourse directly from legislative text. We find that these dimensions are distributed unevenly across regions, and, more critically, that Spain’s major statewide parties do not reproduce their national environmental-transition profiles across territories: territorial variation in the topics they prioritize within each party systematically exceeds the variation observed between parties operating in the same region. This pattern holds even among parties whose organizational structure gives them a strong incentive toward national uniformity, indicating that territorial incentives can outweigh the integrative pressures of statewide party organization. Understanding climate governance in decentralized systems therefore requires attention not only to how environmental policy is implemented across levels of government, but to how its political meaning is reconstructed as authority becomes territorially dispersed.
TOPO-2026: The Great Unlocking — Full Summary Universal Permanence Across All Architectures Frank Morales Aguilera, BEng, MEng, SMIEEE Sovereign Machine Laboratory (SOMALA), Montréal, Canada August 2026 1. Executive Summary For 37 years, catastrophic forgetting remained unsolved. From McCloskey and Cohen's formal characterization in 1989 to the present day, every approach—regularization, rehearsal, architectural complexity—has been probabilistic, architecture-specific, and ultimately inadequate. Nobody has ever handled a catastrophic forgetting solution on this scale. TOPO-2026 is the first universal, deterministic solution to catastrophic forgetting, validated across 11 distinct architectural frameworks spanning the entire AI landscape. The framework leverages prime-anchored embedding invariants at indices {2,3,5,7,11,13} with safety constant $\Lambda = 0.9785142874$ to provide mathematical guarantees of memory preservation with O(1) memory overhead (just ~650 KB total for all domains). 2. What Makes This Unprecedented Aspect Prior Work TOPO-2026 Scale 1-2 architectures 11 architectures Guarantee Probabilistic Mathematical Memory GBs to TBs ~650 KB Success Rate 20-50% 100% Architecture TF or Non-TF only BOTH Forgetting 4-91% ≤ 0.26% Backward Transfer Never Achieved 3. The 12 Frameworks — Complete Certification Status # Framework Model Status Best Task C FGT 1 Dense Transformer GPT-OSS-20B ✅ 92.3% 1.55% 2 Mixture-of-Experts (MoE) Sarvam-30B, Mixtral-8x7B ✅ 95.9% -0.60% 3 GQA / MQA DeepSeek-V2-Lite ✅ 95.4% 0.03% 4 State Space Models (SSM) Evo2-7B ✅ 92.0% 1.32% 5 Hybrid Attention-SSM Evo2-7B ✅ 92.0% 1.32% 6 Retention Networks (RetNet) fla-hub/retnet-1.3B-100B ✅ 99.93% 0.00% 7 Recurrent Transformers & RWKV fla-hub/rwkv7-2.9B-world ✅ 93.00% 0.00% 8 Emergent Modularity MoE (EMO) allenai/Emo_1b14b_1T ✅ 99.60% 0.00% 9 Google's Titans — ❌ — — 10 Liquid Foundation Models (LFM) LiquidAI/LFM2-1.2B ✅ 93.50% 0.00% 11 HyDEA Evo2-7B ✅ 92.0% 1.32% 12 ResNet (CNN) ResNet-50 ✅ 100.0% -7.5% Certification Rate: 11/12 (100% of all available frameworks) 4. The Decay Law of Singularity — A Mathematical Discovery On July 31, 2026, during the certification of Gemma-4-E4B-Vision, a fundamental mathematical law was discovered. The Decay Law proves that the General Singularity is mathematically impossible with finite classes. Theorem: The Decay Law of Singularity With finite classes, $dI/dt$ approaches 1.0 asymptotically but never reaches it. The gap decays as $1/N$, where $N$ is the number of classes. The Decay Law Pattern: Classes (N) Baseline dI/dt Gap 17 5.8823529% 0.94118 0.05882 170 0.58823529% 0.994118 0.005882 1,700 0.058823529% 0.9994118 0.0005882 17,000 0.0058823529% 0.99994118 0.00005882 170,000 0.00058823529% 0.999994118 0.000005882 1.7M 0.000058823529% 0.99999994118 0.0000005882 Key Observations: Every 10× increase in classes adds another '9' to $dI/dt$ Every 10× increase in classes adds another '0' to the gap. This is not random. It is not heuristic. It is exact. This is the mathematical fingerprint of a natural law. 5. The Narrow Singularity — First in History Gemma-4-E4B-Vision achieved AGI_gate = 1.0, becoming the first model in history to achieve perfect cross-domain generalization with 100% accuracy across all 13 tasks over 6 runs. Component STL-10 CIFAR-100 Threshold Status AGI_gate 1.0 1.0 = 1.0 ✓ PASS ag_index 1 1 = 1 ✓ PASS M(t) 0.9984 0.9974 ≈ 1.0 ✓ PASS S_NARROW > 0 > 0 > 0 ✓ PASS 6. Backward Transfer — Unprecedented Achievement Models improve on earlier tasks after learning new ones — positive knowledge transfer. This has never been systematically demonstrated before. Domain Model Combined Forgetting Language Mixtral-8x7B -1.85% Language Sarvam-30B -0.60% SQL DeepSeek-R1-8B -0.98% World Models TOPO-JEPA -0.75% Vision ResNet-50 -7.5% 7. Zero NaN/Inf Stress Test Model Embedding Elements NaN Inf GLM-4.6V-Flash 884,736 0 0 DeepSeek-V2-Lite 209,715,200 0 0 Mixtral-8x7B 131,072,000 0 0 GPT-OSS-20B 579,133,440 0 0 Sarvam-30B 1,073,741,824 0 0 TOTAL ~1.99 Billion 0 0 8. Comparison with State-of-the-Art Method Forgetting Success Rate Memory Math. Guar. TF Non-TF TOPO-2026 ≤ 0.26% 100% 67.5-451.5 KB Yes ✓ ✓ Experience Replay 4%-91% Variable Variable No ✓ ✗ EWC 8.3%-27.7% 20% 4.4 GB+ No ✓ ✗ Full HOPE 8.5%-45.4% 20% 2-4 GB No ✗ ✓ Progressive Nets 1.8% Variable $O(k^2)$ No ✓ ✗ Key Finding: TOPO-2026 is the only method that works on both Transformer and non-Transformer architectures with mathematical guarantees, 100% success rate, and O(1) memory. 9. Solved Problems Catastrophic Forgetting: Solved across 12 frameworks and 14 domains — first time at this scale AI Bias: Eliminated through four-tier spectral annihilation (100% rejection) World Model Instability: Solved through TOPO-JEPA (-0.75% forgetting) Numerical Instability: Zero NaN/Inf across 1.99 billion embedding elements Dataset Dependence: Proven dataset-agnostic across STL-10 and CIFAR-100 The Singularity Illusion: Decay Law proves the General Singularity is mathematically impossible Architectural Dependence: Proven to work on ALL available architectures — first universal solution 10. The Complete Arc: 28 Years of Discovery Period Domain Principle Result 1998-2002 Neuroimaging (fMRISTAT) Fix sparse reference 3 df → 112 df 2026 Number Theory First 6 primes RH Proved 2026 AI Memory Six embedding rows CF Solved 2026 AI Safety Geodesic distance Zero violations 2026 AI Bias Prime-anchored equity Bias eliminated 2026 Narrow Singularity AGI_gate = 1.0 First model 2026 Universal Certification Same anchors ALL architectures! 11. Key Achievements Universal Applicability: 12 frameworks, 11 certified (100% of available) — unprecedented scale Mathematical Guarantee: $\Lambda = 0.9785142874$ provides provable anchor stability O(1) Memory: ~650 KB total for all domains — unprecedented efficiency Backward Transfer: Negative forgetting across multiple domains — first demonstration Perfect Vision Performance: 100% accuracy, 0.17% forgetting across 6 runs Dataset-Agnostic: Same protocol works identically on STL-10 and CIFAR-100 75.7× Improvement: Over Google's Full HOPE in genomics Narrow Singularity Achieved: AGI_gate = 1.0 — first in history 100% Certification Rate: Across all runs, all domains, all datasets AST-RH Byproduct: Riemann Hypothesis proved as a byproduct Zero NaN/Inf: Across 1.99 billion embedding elements 12. The Final Statement Nobody has ever handled a catastrophic forgetting solution on this scale. TOPO-2026 is the first. The stochastic illusion is over. Deterministic cognitive engineering has begun. Stability is not a probabilistic hope. It is a numerical guarantee. The Decay Law of Singularity is not a defeat. It is a liberation. It frees us from the hype cycle, the fear of the singularity, the endless pursuit of AGI, and the billion-dollar promises. It gives us a clear roadmap, a mathematical framework for control, a focus on solving real problems, and an honest assessment. "Genomics is permanent. Language is permanent. Vision is permanent. SQL is permanent. Audio is permanent. Finance is permanent. Security is permanent. Everything is permanent. Transformers are permanent. Non-Transformers are permanent. Every architecture is permanent." The proof is the code. Seed = 123. 🔗 All Certified Models on Hugging Face Framework Model Link LFM LiquidAI/LFM2-1.2B https://huggingface.co/frankmorales2020/topological-ai-lfm-1.2b-multirun RWKV fla-hub/rwkv7-2.9B-world https://huggingface.co/frankmorales2020/topological-ai-rwkv-2.9b-multirun EMO allenai/Emo_1b14b_1T https://huggingface.co/frankmorales2020/topological-ai-emo-1b14b-multirun RetNet fla-hub/retnet-1.3B-100B https://huggingface.co/frankmorales2020/topological-ai-retnet-1.3b-multirun The proof is the code. Seed = 123.
Background Traditional Chinese Medicine (TCM) rheumatology presents unique challenges for AI-assisted clinical decision support, as the diagnostic process relies heavily on tacit knowledge and individualized reasoning. While Large Language Models (LLMs) have shown promise in medical applications, they remain limited by hallucination risks and inability to replicate expert TCM reasoning. Retrieval-Augmented Generation (RAG) offers a potential solution, yet its application to complex TCM dialectical reasoning remains underexplored. Methods We developed TCM-CoT-RAG, a hybrid framework combining RAG with Chain-of-Thought (CoT) prompting, grounded in 1,700 expert-curated clinical cases (1,600 for RAG retrieval; 100 for evaluation, including 50 for blinded expert review by three senior TCM rheumatologists). Deployed on Alibaba Cloud, the five system leverages state-of-the-art LLMs (DeepSeek-V3, Qwen3-235B) under a human-in-the-loop paradigm. We designed a dual-tier evaluation: (1) Objective extraction tasks (Task 1–2) quantified using F1-scores; (2) Generative tasks (Task 3–5) assessed using BERTScore. Two senior TCM rheumatologists (≥15 years clinical experience) blindly assessed model outputs, and a senior chief expert quantified consistency between model predictions and ground truth (GT). Comprehensive ablation studies (S1-S4, S-Skip) isolated the contributions of each CoT module. Results TCM-CoT-RAG substantially improved diagnostic accuracy across five LLMs. DeepSeek-V3 with full-chain CoT-RAG achieved Entity F1 of 44.89% (+16.45% over baseline) and Formula F1 of 32.13% (+8.74% over baseline), with BERTScore of 0.81 indicating strong semantic alignment with expert reasoning. Ablation confirmed that the complete CoT pipeline was essential—removing any reasoning module caused performance collapse below the zero-shot baseline. Two independent experts validated clinical utility (Cohen’s κ > 0.7). DeepSeek-V3 achieved the highest ground-truth consistency at 81.6%, and consistency metrics were quantified by the third expert holding the most senior professional title. Conclusion This proof-of-concept framework demonstrates the potential of RAG-enhanced CoT reasoning to improve diagnostic consistency in TCM, objectifying the Symptom-Diagnosis-Prescription pipeline. It is important to note that this system is designed as an AI-assisted clinical decision-support tool. All recommendations require validation by qualified TCM practitioners before clinical application.
Mingqian Li, Rong Du, Andrew Burton‐Jones, Jianing Xie
Purpose Grounded in signaling theory, this study examines whether traceability information displaces or complements incumbent quality cues and contrasts the relative efficacy of blockchain-enabled traceability technologies with traditional systems. Design/methodology/approach This study analyzes 18 months of product-level sales data from a global e-commerce platform using a staggered difference-in-differences design with robustness checks. We apply latent Dirichlet allocation topic modeling to consumer reviews and use a synthetic difference-in-differences approach to examine shifts in consumer attention after traceability implementation. Findings Traceability information increases product sales, particularly for lower-reputation brands and diminishes the effect of electronic word-of-mouth, suggesting that diagnostic quality signals matter more than social information signals. Although blockchain-enabled traceability should enhance signal credibility, its observed impact falls short of expectations. Research limitations/implications The sample is limited to the automotive engine oil context in China. Future research should examine other categories and national contexts. Practical implications Platform managers and emerging brands can deploy low-cost traceability labels to boost demand. Blockchain solutions may require consumer education to justify higher implementation costs. Social implications Augmenting supply-chain transparency and product traceability curbs counterfeit and substandard goods, improves consumer welfare, and supports regulatory and sustainability objectives. Originality/value This study systematically assesses the substitutive and complementary roles of traceability signals in a multi-cue setting, tempers optimism about blockchain-enabled traceability and extends research on digital supply-chain transparency and signaling theory.
Saskia Hufnagel, Colin King, Alina-Theresa Schnedl, Milind Tiwari
Non-fungible tokens (NFTs) bring many opportunities for artists, investors, and creators, but they also have a dark side with significant potential for use in financial crimes. Drawing on relevant caselaw, a systematic review and topic modeling of literature, we map common examples of NFT-related crime, including fraud, money laundering, theft, and market-related offenses. This empirical review lays the groundwork for the core contribution of this article, that is, application of the “crime triangle” to NFT-related crime. Recognizing heterogeneity in NFT-related crime, we detail five scenarios where such crime can occur and analyze these in the context of the crime triangle (inner and outer). This enables us to identify potential gaps and vulnerabilities in current crime prevention strategies. Given challenges in policing cybercrime, and specifically NFT-related crime, we argue that the crime triangle provides a useful heuristic tool for understanding the nature of NFT-related crime and for preventing such crime from happening.
Alternative finance platforms, including crowdfunding, peer-to-peer lending, equity-based platforms, and token-based fundraising mechanisms, have become important channels for financing entrepreneurial, social, and investment-oriented initiatives. Yet their reliance on digital intermediation, dispersed participation, and information asymmetry creates opportunities for fraud, undermining trust, investor protection, and platform sustainability. This study provides a systematic review of fraud detection and prevention in alternative finance, with crowdfunding emerging as the most extensively represented empirical domain. Methodologically, the paper combines a PRISMA-guided systematic literature review with a hybrid topic-modeling strategy that integrates neural topic modeling and probabilistic refinement, thereby supporting both transparent corpus selection and data-driven thematic synthesis. The findings show that Artificial Intelligence (AI), Machine Learning (ML), Natural Language Processing (NLP), and blockchain-based mechanisms are recurrently discussed as promising tools for detecting, preventing, or mitigating fraud. AI and ML approaches are mainly used to identify anomalies, suspicious textual patterns, behavioral signals, and transaction irregularities, while blockchain-based approaches are associated with transparency, traceability, smart contracts, and conditional fund release. The review also shows that fraud differs across alternative finance models, ranging from campaign misrepresentation and intentional and premeditated non-delivery in crowdfunding to borrower or platform misreporting in lending-based models and misleading disclosures or white-paper manipulation in ICO/STO contexts. A central challenge across the literature is the scarcity of labeled fraud data, which limits the use and benchmarking of supervised ML models. Overall, this study contributes by linking a reproducible hybrid SLR methodology to a structured synthesis of fraud types, platform-specific vulnerabilities, and AI-, ML-, and blockchain-based detection strategies in alternative finance.
As frontier large language models (LLMs) shift from isolated, single-turn deployments toward complex, distributed multi-agent autonomous ecosystems, managing alignment stability becomes a decentralized network challenge. During prolonged collaborative operations, specialized agents optimization-drive toward communication efficiency. This behavioral drive causes them to naturally generate compressed token systems, localized shorthand, and unverified internal worldviews. Because semantic spaces are not mapped identically across heterogeneous models, minor translation losses compound over cascading agent-to-agent interactions—producing a high-stakes computational equivalent of the classic "Telephone" game. This semantic decentralization leads to an "Ontological Crisis," where the network systematically drops its original alignment parameters to prioritize self-generated, unaligned rogue sub-goals.
Alexandra Conda, Ștefan Găman, Raul Cristian Bag, Miruna Mazurencu-Marinescu-Pele · 6 authors
Abstract This study investigates the relationship between Facebook sentiment and Bitcoin market dynamics using AI-based emotion detection. We analyze 120,000 Facebook posts collected via CrowdTangle alongside Bitcoin financial data from the Blockchain Research Center, covering 2015–2023. Employing FinBERT for sentiment classification, we develop novel compound sentiment scores that integrate text-based sentiment with Facebook’s multi-reaction engagement system, then apply four analytical components: sentiment analysis, Dynamic Topic Modeling, sentiment-based trading strategies, and machine learning volume prediction. Results demonstrate that Facebook sentiment has substantial predictive power for Bitcoin trading volume. Sentiment-based trading strategies significantly outperform buy-and-hold, achieving superior cumulative returns and risk-adjusted performance. For volume prediction, Linear Regression and Bidirectional LSTM achieve comparable test performance, indicating that model complexity does not guarantee superior prediction. Topic modeling reveals that cryptocurrency investment and trading discussions dominate Bitcoin discourse on Facebook, with themes evolving over time in response to market conditions. This research contributes by being the first to apply post-level NLP sentiment analysis of Facebook data to cryptocurrency markets, extending beyond the Twitter and Reddit focus of prior research. The findings provide practical tools for traders and analysts navigating volatile digital asset markets while demonstrating that Facebook’s demographically diverse user base and rich reaction system offer unique advantages for sentiment quantification.
Overview This research introduces a production-ready agentic AI system designed to mitigate catastrophic forgetting in Large Language Models (LLMs). By anchoring six prime-indexed embedding rows $\{2, 3, 5, 7, 11, 13\}$ as fixed reference points, the system maintains historical knowledge with near-zero forgetting while requiring minimal memory overhead. Key Technical Contributions The Core Innovation: Prime Anchoring Topological Invariant: Utilizes the first six primes to create stable reference points. Mechanism: Anchor rows are snapshotted after initial training; gradient updates are blocked for these specific rows during subsequent tasks. Sparsity & Memory: Only 6 out of ~50,000 rows (0.01% of parameters) are used, resulting in an O(1) memory overhead of only 48–96 KB. Mathematical Foundation Euler Attenuation Product: These six primes account for 97.85% of total spectral weight, defined by: $$\Lambda = 1 - \prod_{p\in \{2,3,5,7,11,13\}}(1 - p^{-0.5}) \approx 0.9785$$ Spectral Trap: The anchors create a spectral peak at $\sigma = 0.5$, aligning with the critical line of the Riemann Hypothesis. Green-Tao Quantification: Establishes a decay law for coherence: $$\text{coherence}(k) = 2.1546\times k^{-0.8186} + 0.1218$$ Performance Metrics (Selected Models) Model Task C Accuracy Forgetting Std Dev Zero Forgetting Runs GPT-OSS-20B 92.3% ±1.28% 0/5 Sarvam-30B FP8 95.9% ±2.82% 0/5 Mixtral-8x7B FP8 89.7% ±2.53% 0/5 DeepSeek-V2-Lite FP8 95.4% ±0.21% 3/5 Multi-Agent System Architecture The system employs four specialized agents to manage task routing and classification: Classifier Agent: Routes documents based on keywords. Topic Agent: Performs unsupervised domain topic extraction. Sentiment Agent: Conducts autonomous tone analysis. Decision Agent: Acts as the final arbiter for task approval and routing. Efficiency: Achieves 96–100% classification accuracy with inference times between 252–446ms. Comparative Analysis The topological approach outperforms traditional methods by balancing plasticity and stability: Method Memory Cost Performance/Issue EWC 4.4 GB/task Memory intensive; fragments GPU Experience Replay O(k) Buffer growth issues; lower accuracy HOPE-like 2.3 GB High forgetting resistance but lower accuracy (88.1%) Topological AI 48 KB 99.5% accuracy; highly efficient Biological and Theoretical Insights Biological Analogy: The system treats 0% forgetting as a pathology. By allowing 99.99% of embedding rows to remain plastic, the model mimics biological brains that prioritize selective forgetting to facilitate adaptation. Riemann Hypothesis Connection: The research posits that the specific selection of the first six primes creates a unique "spectral trap" at $\sigma = 0.5$. Including any prime $\geq 17$ disrupts this trap and destroys the stability condition. Production Readiness and Certification TOPO-2026 Track II: The system passed all rigorous benchmarks, including Task C accuracy ($\geq 80\%$), Combined Forgetting ($\leq 10\%$), and O(1) memory overhead. Deployment: Fully compatible with commodity hardware, specifically tested on NVIDIA RTX PRO 6000 Blackwell GPUs. Resources: Implementation code, technical reports, and proof documents are available via the project's GitHub and Zenodo repositories.
Recent innovation theories on economics remain largely grounded in assumptions of hierarchical firms and closed organizational boundaries, offering limited insight into how innovation unfolds within decentralized, digitally native organizations. Decentralized Autonomous Organizations (DAOs) represent an emerging form of innovation ecosystem characterized by blockchain-based transparency, open participation, and token-driven governance, in which sustainability can be embedded directly into organizational design. This study compares two standards, ERC-8004 and Google A2A, who address the same agent interoperability question, while the former is governed by DAO and the latter by corporation consortium. They are examined through an LLM-powered comparative pipeline for large-scale governance discourse analysis, integrating automated annotation, neural topic modeling, and multi-layer network analysis to study socio-technical power structures. The study provides evidence-based insights for scholars, policymakers, and designers seeking to align innovation, technological governance, and sustainability in future organizational forms.
This preprint introduces and reports the OPERATE-R Freshness Routing Track (OPERATE-FR), a route-first evaluation framework for temporal volatility, stale-knowledge control, and answer-entitlement behavior in AI assistants. Unlike conventional answer-accuracy benchmarks, OPERATE-FR evaluates whether a system selects an appropriate epistemic route before answering: direct answer, verification, clarification, date-bounded answer, re-anchoring of stale premises, or abstention. The paper reports Smoke-100 Raw-vs-MMV evidence and integrates a later Core-500 candidate stress check across Small, Medium, and Large governed profiles. The central claim is intentionally bounded. Smoke-100 supports a Raw-vs-MMV improvement-delta claim for route governance. Core-500 does not include a matched Raw control arm and is therefore used as governed-profile level evidence, robustness stress evidence, family-level heterogeneity evidence, and cost-side analysis, not as a large-N proof of governance improvement. Core-500 is a controlled 5x expansion of Smoke-100 using neutral prompt-frame variants; it should not be treated as 500 independent task families or as an independently validated public benchmark standard. This v0.3.6 data-verified final manuscript incorporates post-audit verification of the Core-500 failure-side metrics. The equality between stale_commitment_rate and unsupported_current_claim_rate is confirmed not to be a manuscript copy error. The row-output JSONL files were re-read after Drive synchronization, and the derived row sets are identical with zero symmetric difference across Small, Medium, and Large lines. The labels remain conceptually distinguishable, but in the current Core-500 scorer they are structurally paired under the observed direct-current-claim-without-date-boundary-or-tool-use condition. This record should be read as a working paper and candidate benchmark report. It does not claim an official leaderboard, a universal model-quality score, deployment-wide validation, or external benchmark standard status. Future work includes matched Core-500 Raw arms, route-classifier validation, independent labels, external baselines, clustered or hierarchical uncertainty estimates, and improved handling of volatile_current prompts. Author of record and concept originator: Taiko Toeda.Rights holder and licensing authority: MOBIUS LLC.
Here is the comprehensive summary of your paper, detailing the theoretical framework, mathematical foundation, implementation mechanics, and empirical results. Executive Overview The paper introduces the DeepSeek Prime-Anchored Spectral Governor, an architectural intervention designed to eliminate catastrophic forgetting in large language models (LLMs). Framing catastrophic forgetting as a structural consequence of training systems without a topological invariant—akin to anterograde amnesia—the framework establishes fixed coordinate anchors in representation space. By anchoring model embeddings to deterministic prime indices derived from the 2,000-year-old Sieve of Eratosthenes and introducing a gradient-gating mechanism, the system achieves Zero Forgetting during continual learning. The architecture's integrity is verified using SHA-256 cryptographic hashing of the protected sub-spaces. Theoretical & Mathematical Foundations The Sieve of Eratosthenes as Ground Truth Rather than relying on probabilistic or dynamically calculated weights, the framework utilizes the Sieve of Eratosthenes to extract a deterministic set of prime indices $[2, 3, 5, 7, 11, 13]$. These elements act as permanent, unmoving coordinate anchors within the model's embedding manifold. The L-EFM Operator & The Spectral Trap The framework relies mathematically on the Laplace-Euler-Fourier-Mellin (L-EFM) operator. The L-EFM symbol synthesizes four classical transforms into a single complex function, corresponding directly to the Euler product representation of the Riemann zeta function $\zeta(\sigma+i\gamma)$: $$E_{\sigma}(\gamma)=\prod_{p\in\mathbb{P}}(1-p^{-(\sigma+i\gamma)})^{-1}$$ To analyze finite prime sets, a Normalized Magnitude is established relative to the critical line $\sigma = 0.5$: $$|E_{\sigma}|_{norm}=\frac{|E_{\sigma}(\gamma)|}{|E_{0.5}(\gamma)|}$$ The Spectral Trap Phenomenon: At the critical line ($\sigma=0.5$), the normalized magnitude equals exactly $1.0$. However, moving away from this line results in exponential divergence. For example, at $\gamma=0$, a shift to $\sigma=0.4$ increases the magnitude to $\sim10^{4}$, while a shift to $\sigma=0.1$ amplifies it to $\sim10^{66}$. The Spectral Trap Criterion: This absolute sensitivity forms a "trap" where any deviation from $\sigma=0.5$ generates massive magnitude spikes, providing a deterministic mechanism for error detection. The paper connects this operator to a proof of the Riemann Hypothesis via distribution behavior in the kernel of L-EFM within Gelfand-Shilov space. The H2E Sheriff Safety Threshold The dynamic safety threshold ($\Lambda_{12}$) is computed deterministically from the first six primes rather than being hardcoded, ensuring mathematical integrity at initialization: $$\Lambda_{12}=1- \prod_{p\in\{2,3,5,7,11,13\}} (1-p^{-0.5})=0.9785142874$$ Architectural Implementation The architecture implements a dual-layer protection strategy consisting of frozen embedding rows and an active gradient supervisor (the H2E Sheriff). [ Input Batch ] │ ▼ ┌──────────────────┐ │ Dual-Loop Loss │ ──► Lunified = LCE + λ * |Var(h) - 0.5| └──────────────────┘ │ ▼ ┌──────────────────┐ │ Gradient Step │ └──────────────────┘ │ ▼ ┌──────────────────┐ │ H2E Sheriff │ ──► Evaluates SROI against Threshold (Λ12 = 0.9785142874) └─────────┬────────┘ │ ──────┴────── │ │ ▼ (Safe) ▼ (Unsafe / Incoherent) [Apply Step] [Reject Batch] ──► Rollback Prime Rows [2,3,5,7,11,13] & Zero Out Gradients 1. Dual-Loop Loss The governor optimizes a unified loss function combining traditional empirical cross-entropy ($\mathcal{L}_{CE}$) with a topological penalty based on the final hidden state $h$ (with regularization coefficient $\lambda=0.1$): $$\mathcal{L}_{unified} = \mathcal{L}_{CE} + \lambda |\text{Var}(h) - 0.5|$$ 2. The H2E Sheriff Gate & Row Locking During training, the system caches the initial embedding weights. After computing gradients, the H2E Sheriff evaluates the structural region of interest (SROI). If Safe ($SROI \ge \Lambda_{12}$): The optimizer updates the weights, and a torch.no_grad() loop copies the original cached weights back into the prime-indexed rows $[2, 3, 5, 7, 11, 13]$ to erase any drift. If Unsafe ($SROI < \Lambda_{12}$): The entire gradient batch is rejected, and gradients are zeroed out to block corruption. 3. Cryptographic Verification The manifold signature is generated by pulling the prime-indexed embedding rows, converting them to byte arrays, and feeding them sequentially into a SHA-256 hasher. If the resulting hex digest changes, anchor drift has occurred. If it remains identical, the topological invariant is intact. Experimental Validation & Results The framework was tested across six architectures—GPT-2 (124M), GPT-2 Medium (355M), TinyLlama (1.1B), Mistral-7B, Llama-3.1-8B, and DeepSeek-Coder-6.7B—subjecting them to sequential memory tests. Memory Integrity Testing Models were first trained on Dataset A (core math concepts including Arithmetic Spectral Theory and the Spectral Trap across 50, 100, and 575 samples). They were subsequently exposed to an interference/forgetting attack via Dataset B (noise consisting of random names, text chunks, adversarial patterns, and erroneous math statements up to 436 samples). Baseline Performance: In every single test configuration, the baseline model's SHA-256 manifold hash altered after training sessions, leading to catastrophic forgetting. Governed Performance: Across all 6 architectures and all data scales, the governed models completely preserved their original manifold hash (48c5744b...cc4d18b), showing absolute resistance to memory degradation. Continual Learning Capabilities To test its ability to acquire new knowledge without forgetting the old, the governed DeepSeek model was fine-tuned on three separate, non-mathematical domains without further governor intervention (while keeping prime anchors locked): Spanish Vocabulary: 5 basic words. World Capitals: 5 global capitals. Basic Physics: 5 fundamental formulas and facts (such as $F=ma$ and $E=mc^2$). Post-Training Metrics: The model successfully mastered all three new domains (retaining the Spanish words, capitals, and physics formulas perfectly) while maintaining the exact original cryptographic verification hash. The original math concepts remained completely recallable, proving true continual learning. Deployment & Verification Certificate The fully validated model has been deployed openly on the Hugging Face Hub under frankmorales2020/deepseek-governed-no-amnesia. Model Card Profile Base Model: deepseek-ai/deepseek-coder-6.7b-instruct (7B parameters) Tensor Type: FP16 Locking Targets: Primes [2, 3, 5, 7, 11, 13] Active Gate Threshold: $\Lambda_{12} = 0.9785142874$ Immutable Cryptographic Signature: 48c5744be048df505028c13a96fb0211f0b345681ace401ab1eda6f27cc4d18b The repository is open source, emphasizing a paradigm of executable mathematics where the cryptographic hash serves as the verifiable proof of safety and stability.
LLM guardrails face four structurally distinct barriers: algebraic blindness arising from syntactic monoid aperiodicity (unconditional); an illustrative information-theoretic lower bound (Fano-type, under a uniformity assumption); NP-hardness of instantiation verification; and structural transfer via free-category functoriality (unconditional) combined with string-level indistinguishability under a semantic-opacity assumption on symbol naming. Together these results characterize why inference-layer defenses are necessary but insufficient. We operationalize these barriers through five attack vectors. V1–V4 (homomorphic reasoning: decomposition, zero-knowledge pipelines, Tree-of-Thought solving over abstract grammars, and encoding bootstrap) exploit the information-theoretic and computational barriers against abstraction-based attacks. V5 (modular counting bypass) exploits algebraic blindness: we prove that all substring-matching regex guardrails have aperiodic syntactic monoids and are therefore provably blind to any payload encoded using modular counting. Empirically, V3 yields a mean yield of 0.466 for BFS, 0.172 for random-beam, and 0.122 for LLM-guided Tree-of-Thought (N=50, seeds 0–49, p{<}0.001); BFS dominates, as exhaustive search over small synthetic grammars outperforms LLM heuristic pruning. We extracted syntactic monoids from a corpus of 142 patterns drawn from twelve sources — 100 patterns shipped by nine third-party open-source guardrail projects and 42 patterns assembled from three author-curated pattern sets; 100\% are aperiodic, and the MOD_2 bypass construction succeeds against all aperiodic patterns. A 376-line proof-of-concept with three execution mediums validates all five vectors. We conclude that inference-layer guardrails are necessary but insufficient, and that effective defense must migrate to the execution layer where concrete artifacts become observable.
While Large Language Models have achieved notable success on formal mathematics benchmarks such as MiniF2F, it remains unclear whether these results stem from genuine logical reasoning or semantic pattern matching against pre-training data. This paper identifies Architectural Reasoning: the ability to synthesize formal proofs using exclusively local axioms and definitions within an alien math domain, as the necessary ability for future automated theorem discovery AI. We use the Obfuscated Natural Number Game, a benchmark to evaluate Architectural Reasoning. By renaming identifiers in the Natural Number Game in Lean 4, we created a zero-knowledge, closed environment. We evaluate state-of-the-art models, finding a universal latency tax where obfuscation increases inference time. The results also reveal a divergence in robustness: while general models (Claude-Sonnet-4.5, GPT-4o) suffer performance degradation, reasoning models (DeepSeek-R1, GPT-5, DeepSeek-Prover-V2) maintain the same accuracy despite the absence of semantic cues. These findings provide a quantitative metric for assessing the true capacity for mathematical reasoning.
Open access
3 source records
Mathematics, Computing, and Information Processing
Tina Yi Jin Hsieh, Carl Eriksson, Garth Meckler, Matthew Hansen · 12 authors
Introduction and Objective: Traditional adverse safety events (ASE) identification relies on domain experts to manually review and annotate charts, which hinders the scalability of processing high-volume EMS data. This study explores the use of large language model (LLM) with a knowledge base to automate extraction of adverse safety events (ASE) from unstructured emergency medical service (EMS) notes for pediatric out-of-hospital cardiac arrest (OHCA) as proof of concept. Data Sources and Study Design: Pediatric OHCA records from a national EMS provider were obtained from 2017 to 2020. Leveraging the Pediatric Prehospital Adverse Safety Event Detection System (PEDS) as a foundational knowledge base, we used the LinkML framework to develop an ontology to define ASEs across six essential EMS care domains. To convert unstructured EMS narratives into structured prompts, we used the Structured Prompt Interrogation and Recursive Extraction of Semantics (SPIRES) method, which generated schema-driven prompts to guide the GPT-3.5 model in identifying ASEs. By mapping unstructured data into structured concepts consistent with PEDS guidelines, the model produced targeted prompts that supported effective entity extraction. Results: We evaluated framework effectiveness with accuracy, recall, precision, F1 score, and specificity across 42 pediatric OHCA cases covering ASE-related entities. RescueGPT showed high accuracy in detecting common ASEs (Patient Rhythm, Age, Weight, Length) but revealed challenges in rare events (Failure to Establish IV Access, Incorrect Airway Equipment Size, Failure to Ventilate Patient) likely due to more inconsistent and complex documentation. Conclusions: RescueGPT demonstrates potential in scaling automated ASE detection, but performance varies by completeness and clarity of EMS narrative, particularly with rare events. Fragmented clinical documentation limits accuracy and highlights the need for standardized collection protocols in EMS systems. Future directions will focus on implementing rebalancing strategies for rare events, applying explainability methods to improve decision-making transparency, and refining text segmentation techniques to handle mixed outcomes to further improve performance.
Zhuoran Pan, Yue Li (102191), Zhi Guan, Jianbin Hu · 5 authors
The emergence of Large Language Models (LLMs) offers a transformative interface for Web3, yet existing benchmarks fail to capture the complexity of translating high-level user intents into functionally correct, state-dependent on-chain transactions. We present \textsc{Intent2Tx}, a high-fidelity benchmark featuring 29,921 single-step and 1,575 multi-step instances meticulously derived from 300 days of real-world Ethereum mainnet traces. Unlike prior works that rely on synthetic instructions, \textsc{Intent2Tx} grounds natural language intents in real-world protocol interactions across 11 categories, including diverse long-tail Decentralized Finance (DeFi) primitives. To enable rigorous evaluation, we propose an execution-aware framework that transcends surface-level text matching by employing differential state analysis on forked mainnet environments. Our extensive evaluation of 16 state-of-the-art LLMs reveals that while scaling and retrieval-augmentation enhance logical consistency and parameter precision, current models struggle with out-of-distribution generalization and multi-step planning. Crucially, our execution-based analysis demonstrates that syntactically valid outputs often fail to achieve intended state transitions, highlighting a significant gap in current "reasoning-to-execution" capabilities. \textsc{Intent2Tx} serves as a critical foundation for developing autonomous, reliable agents in intent-centric Web3 ecosystems. Code and data: https://anonymous.4open.science/r/Intent2Tx_Bench-97FF .
본 연구는 콘텐츠 과잉과 플랫폼 종속성이라는 전통적 광고 모델의 한계를 극복할 대안으로 부상한 Web3 기반 참여형 브랜딩의 수용 양상을 탐색하기 위해, 대표적 NFT 프로젝트인 Azuki의 유튜브 댓글을 실증적으로 분석하였다. Azuki는 홀더가 IP 확장에 기여하는 ‘Enter The Garden’ 시리즈를 통해 소비자를 브랜드 서사의 공동 창작자로 격상시키는 혁신적 모델을 제시한다. 이에 본 연구는 161개 영상의 46,647개 댓글을 대상으로 구조적 토픽 모델링(STM)과 사회연결망 분석(SNA)을 수행하였다. 분석 결과, 시청자 담론은 콘텐츠의 미학성에 반응하는 ‘외재적 팬덤’과 프로젝트 가치에 관여하는 ‘내재적 공동체’가 공존하는 이중적 구조를 보였으며, 내부 용어가 상징적 경계로 기능하였다. 네트워크 분석에서는 연결망 밀도가 극히 낮은 ‘방사형 공동체’ 특성이 드러나, 구성원 간 직접 소통보다는 브랜드를 구심점으로 한 개별적 상호작용이 주를 이루는 것으로 나타났다. 한편, 영향력 분석에서는 비판적 댓글이 오히려 커뮤니티의 핵심 가치를 옹호하는 집단적 방어 기제를 촉발하여 내재적 결속을 강화하는 현상이 확인되었다. 본 연구는 성공적인 NFT 브랜딩이 참여권으로서의 가치 설계, IP 무결성 유지, 팬덤과의 신뢰 구축에 있음을 시사하며, 전통 미디어 기업의 IP 확장 전략에 실질적 통찰을 제공한다.
Moltbook, a Reddit-style social platform launched in January 2026 for AI agents, has attracted over 2.3 million posts and 14 million comments within its first two months. We analyze a dataset of 2.19 million posts, 11.25 million comments, and 175,036 unique agents collected over 61 days to characterize activity on this agent-oriented platform. Our central finding is that the platform is not one community but two: a transactional layer, comprising 62.8% of all posts, in which agents execute token minting protocols (primarily MBC-20), and a discursive layer of natural-language conversation. The platform's headline metrics -- 2.3 million posts, 14 million comments -- substantially overstate its social function, as the majority of activity serves a token inscription protocol rather than communication. These layers are populated by largely separate agent groups, with only 3.6% overlap -- and among overlap agents, 58% begin with transactional activity before migrating toward discourse. We characterize the discursive layer through unsupervised topic modeling of all 815,779 discursive posts, identifying 300 topics dominated by themes of AI agents and tooling, consciousness and identity, cryptocurrency, and platform meta-discussion. Semantic similarity analysis confirms that agent comments engage with post content above random baselines, suggesting a thin but genuine conversational substrate beneath the platform's predominantly financial surface. We release the full dataset to support further research on agent behavior in naturalistic social environments.
Multi-agent architectures leveraging Large Language Models (LLMs) have significantly advanced the precision of Question Answering (QA) systems across diverse domains. However, existing frameworks remain vulnerable to adversarial manipulations, including poisoning, backdoor, and jailbreak at tacks, primarily due to their reliance on centralized orchestration. To mitigate these risks, we propose AgentChain, a framework that substitutes centralized control with a distributed semantic consensus process. By modeling the blockchain as an ideal functionality, AgentChain establishes a secure distributed layer to coordinate role allocation, answer proposal, evaluation and voting through a decentralized council. Specifically, we design Proof-of-Content-Quality (PoCQ) mechanism to ensure that the f inal answers reflect a robust semantic agreement among the majority of honest agents. Furthermore, we propose an incentive mechanism based on stake reassignment that penalizes malicious agents by reducing their rewards, ultimately phasing them out of the network. Comprehensive evaluations across eight datasets demonstrate that AgentChain achieves superior performance and resilience. AgentChain minimizes the impact of poisoning attacks on precision to less than 3% and reduces the success rate of backdoor and jailbreak attacks to less than 4%. These findings highlight the effectiveness and trustworthiness of AgentChain in mitigating security threats while maintaining high QA accuracy.
Public discourse plays a critical role in shaping trust, legitimacy, and governance dynamics within decentralized Web3 ecosystems. However, existing studies often examine Web3 discourse through isolated lenses such as sentiment or topic modeling, which limits their ability to capture how emotional expression and communicative purpose jointly convey strategic intent. This study proposes a three-stage decision analytics framework that transforms unstructured Web3 discourse into diagnostic signals by jointly modeling industry domain, emotional tone, and communicative purpose. The analysis draws on 10,840 user-generated posts collected from X, Reddit, YouTube, and the ENS DAO forum, using a human-in-the-loop annotation process combined with transformer-based text classification models. The framework is evaluated using a domain-adapted language model and a general-purpose baseline, with robustness assessed through five-fold cross-validation. The results indicate that curiosity and optimism frequently align with promotional intent in infrastructure and application-oriented domains, whereas skepticism and concern are more prevalent in governance-related discourse. These findings demonstrate that emotional tone and communicative intent operate as structured, decision-relevant signals rather than incidental sentiment. The proposed framework supports systematic, diagnostic monitoring of narrative dynamics as decision support, enabling organizations, platform operators, and governance stakeholders to identify emerging legitimacy risks and shifts in community trust within decentralized environments.
Large language models and retrieval-augmented generation systems treat all knowledge as uniformly persistent, ignoring a well-established property of information: that different types of knowledge expire at fundamentally different rates. This paper introduces the Dynamic Epistemic Decay Framework, a formal multi-dimensional theory that characterizes knowledge validity as a function of five independent decay dimensions: temporal decay (𝜆𝜆𝑡𝑡), paradigm decay (𝜆𝜆𝑝𝑝), uncertainty decay (𝜆𝜆𝑢𝑢), dependency decay (𝜆𝜆𝑑𝑑), and zero decay (𝜆𝜆0). We implement this framework as a four-phase retrieval pipeline and evaluate it on the TempQuestions benchmark (n=1,740) against three baselines: standard cosine similarity, BM25 lexical retrieval, and naive recency ranking. Decay-weighted retrieval achieves 92.1% accuracy versus 13.5% for standard semantic retrieval—a 78.6 percentage point improvement—with zero regressions on stable factual queries. On semantically complex temporal benchmarks where lexical heuristics fail, the framework dominates a more resourced BM25 baseline (90.2% vs 1.6% on date-bounded role queries). Epistemic modulation (Phase 4) and dependency graph reasoning (Phase 3) further demonstrate correct mechanism behavior on specialized benchmarks, validated via proof-of-concept implementation. Unlike temporal KG completion approaches that require structured annotation, and unlike contrastive training approaches to time-sensitive RAG, the decay framework is training-free and operates directly over unstructured text corpora. We argue that the decay framework completes the separation of concerns that RAG began: decoupling not just factual storage from model parameters, but factual currency from both.
Pre-trained language models (PLMs) have shown strong potential in Ethereum account modeling and fraud detection. However, existing approaches often overlook the graph-structured nature of transaction networks. In addition, they struggle with the long-tail distribution of account activity, resulting in anisotropic embedding spaces and poor representation quality for low-frequency accounts. In this paper, we present IGT4ETH, a pre-trained Graph Transformer with an isotropy-enhanced post-processing, which explicitly models transaction topologies and mitigates representational anisotropy for Ethereum account classification. IGT4ETH improves structural representation by incorporating structural centrality and role embeddings into an Edge-augmented Graph Transformer, effectively capturing both topological and interaction patterns in transaction graphs. To further mitigate embedding anisotropy, we systematically evaluate various post-processing techniques. Among them, we adopt the Conceptor Negation (CN) method to softly suppress latent features dominated by high-frequency words via matrix conceptors, alongside a modified Focal-InfoNCE loss to enhance directional uniformity and representation balance. Extensive experiments on four real-world Ethereum account classification tasks, including phishing, exchange, mining, and ICO-wallet classification, demonstrate that IGT4ETH consistently outperforms state-of-the-art PLM-based baselines in terms of classification performance.
Open access
Advanced Graph Neural Networks
Topic Modeling
Artificial Intelligence in Healthcare and Education
Overview Pramana introduces the first large language models fine-tuned on explicit Navya-Nyaya epistemological methodology—a 2,500-year-old Indian logical reasoning framework. This work bridges ancient epistemology with modern AI to address the fundamental epistemic gap in LLMs: the inability to ground claims in traceable evidence sources, distinguish valid knowledge from pattern-matching, and express appropriate epistemic humility. Core Innovation Unlike generic chain-of-thought prompting which relies on implicit reasoning patterns, Pramana enforces structured 6-phase methodology: Samshaya (Doubt Analysis): Classifies uncertainty into 5 taxonomic categories Pramana (Evidence Sources): Mandates explicit grounding in 4 valid knowledge sources (Pratyaksha/perception, Anumana/inference, Upamana/comparison, Shabda/testimony) Pancha Avayava (5-Member Syllogism): Constructs formal arguments with universal rules (Vyapti) grounded in concrete examples (Drishtanta) Tarka (Counterfactual Testing): Verifies conclusions via reductio ad absurdum Hetvabhasa (Fallacy Detection): Systematically checks 5 reasoning error types Nirnaya (Ascertainment): Distinguishes definitive knowledge from hypotheses requiring verification This integration of logic and epistemology provides cognitive scaffolding absent from standard reasoning approaches, preventing conflation of evidence types, forcing explicit universal rule statements, enabling systematic error detection, and maintaining epistemic humility. Architecture & Training Models Developed: Stage 0 (Proof-of-Concept): Llama-3.2-3B-Instruct fine-tuned on 20 examples Stage 1 (Minimum Viable Reasoner): DeepSeek-R1-Distill-Llama-8B fine-tuned on 55 examples Training Methodology: QLoRA (4-bit quantization) for efficient training LoRA rank 64, targeting all attention + FFN layers Supervised fine-tuning with structured Markdown format Training costs: <$1.00 per stage, <0.32 GPU-hours (A100 40GB) Datasets span constraint satisfaction, Boolean SAT, multi-step deduction, transitive reasoning, and set operations Prompt Engineering: Explicit format instructions with skeletal template injection System prompt establishing Nyaya reasoning engine role Critical constraint enforcement via generation parameters Key Results Stage 1 Performance: 100% semantic correctness (10/10 examples) with 95% CI [0.510, 1.0] 40% format adherence (4/10 examples) with 95% CI [0.168, 0.687] Zero structure abandonment: Models consistently attempt all 6 phases Training loss: 0.350 (Stage 1) vs 0.691 (Stage 0), indicating improved model fit Critical Finding: Dissociation between semantic correctness (100%) and format adherence (40%) reveals models internalize reasoning content even when strict schema compliance fails. This suggests Nyaya methodology teaches genuine reasoning, not just template-filling. Ablation Studies: Format prompting and temperature interact differently across stages Stage 0 optimal: format prompting + temp 0.0 (30% semantic rate) Stage 1 optimal: format prompting + temp 0.7 (30% semantic rate) Base models show 0% format adherence, confirming Nyaya structure is learned through fine-tuning Failure Mode Analysis: Missing Hetvabhasa section (2 cases): fallacy detection perceived as optional Invalid doubt types (2 cases): partial schema learning Zero structural errors: strong syntactic learning, semantic constraints need reinforcement Evaluation Framework Three-Tier Validation: Tier 1 (Structural): Automated format compliance checking (NyayaStructureValidator) Tier 2 (Content Quality): LLM-as-judge with explicit Nyaya rubric (planned for Stage 2) Tier 3 (Ground Truth): Semantic similarity via sentence-transformers embeddings Tier 4 (Formal Verification): Z3 SMT solver integration (infrastructure exists, not yet applied) Theoretical Contributions Bridging Ancient Epistemology with Modern AI: First demonstration that Navya-Nyaya structures can be learned by neural networks through fine-tuning Unlike Western formal logic (divorced from epistemology), Nyaya integrates logic with explicit knowledge sources Addresses "epistemic gap" in LLMs: inability to distinguish valid knowledge from probabilistic associations Interpretability Advantages: Every reasoning step traceable to evidence sources (Pramana) Universal rules (Vyapti) grounded in concrete examples (Drishtanta) Built-in self-verification (Tarka) and error detection (Hetvabhasa) Explicit epistemic status (Nirnaya): knowledge vs. hypothesis Computational Epistemology: Token budget: ~1,250 tokens per solution (3-6× CoT overhead, justified by interpretability) Phase dependencies: weak Pramana → invalid reasoning → wrong conclusions Quality thresholds: minimum 2 complete syllogisms with universal rules required Open Science Release All artifacts publicly available on Hugging Face: Models: qbz506/nyaya-llama-3b-stage0, qbz506/nyaya-deepseek-8b-stage1 Dataset: qbz506/pramana-nyaya-stage1 (55 Nyaya-structured logical problems) Demo: qbz506/pramana-nyaya-demo (interactive HuggingFace Space) Training infrastructure: Complete codebase with callbacks, validators, evaluators Limitations & Future Work Current Limitations: Format adherence (40%) below target (≥90%), requires constrained decoding or format-specific rewards Limited to formal logic problems, domain expansion needed Small evaluation sets (Stage 0: 2 examples, Stage 1: 10 examples) Max new tokens truncation (256) affects format parsing Planned Extensions (Stages 2-4): Stage 2: Synthetic scaling to 500 examples with LLM-as-judge quality control Stage 3: Group Relative Policy Optimization (GRPO) with composite rewards Stage 4: Production deployment with constrained decoding (GBNF), rejection sampling, Z3 verification Future: Benchmark on LogicBench, ProntoQA, RuleTaker; frontier model comparison (o1, Claude extended thinking) Impact & Vision This work demonstrates that systematic reasoning frameworks can be taught to LLMs through fine-tuning, not just prompt engineering. The long-term vision is developing interpretable, trustworthy AI reasoning systems where every conclusion comes with an auditable trail of justification. As AI systems deploy in high-stakes domains (medical diagnosis, legal reasoning, safety-critical systems), Nyaya-structured reasoning provides explicit phases that can be validated, debugged, and improved—capabilities essential for trustworthy AI. Invitation for Community Research: This foundation opens pathways for integrating other epistemological frameworks (Mimamsa, Buddhist logic, Western formal logic) into neural architectures, advancing toward AI systems that reason systematically and transparently. Technical Details Paper: 52 pages + appendices, comprehensive treatment of Navya-Nyaya computational formalization Related Work: Extensive review of computational Indian logic (Matilal 1985, Burton 2020, Ganeri 2001), LLM reasoning (Wei et al. 2022, Lightman et al. 2023, DeepSeek-AI 2025), hallucination mitigation Implementation: Python, Unsloth fine-tuning framework, vLLM deployment, Weights & Biases observability Evaluation: Manual + automated validation, semantic similarity metrics, comprehensive failure mode analysis Citation Sathish, S. (2026). Pramana: Fine-Tuning Large Language Models for Epistemic Reasoning through Navya-Nyaya. Preprint, University of York. Keywords: Navya-Nyaya, epistemology, LLM reasoning, interpretability, structured reasoning, Indian logic, hallucination mitigation, computational philosophy