Blockchain Papers

Follow blockchain research across journals, conferences, and preprint repositories.

89 papersLast indexed Aug 31, 2026
Search papers

Paper index

89 results · page 1 of 4

Clear filters
Aug 22, 2026·Zenodo (CERN European Organization for Nuclear Research)
0 cites
FCD-TITAN 2-C: Finite-Frame Transfer Compatibility: Exact Spectral-Mixing Residuals under Known Circular LTI Transfer

Geoffrey Marcellin

Finite-Frame Transfer Compatibility: Exact Spectral-Mixing Residuals under Known Circular LTI Transfer Finite spectral estimation and known linear transfer do not generally commute. This work shows that the resulting spectral discrepancy is not merely an uncontrolled finite-window artifact: under a known circular LTI transfer and an exact all-origin finite-frame mixing kernel, the compatibility residual is analytically defined and quantitatively computable without fitted calibration. For an input power spectrum S, a nonnegative finite-frame mixing operator K, and power transfer a=∣H∣2, the fitted log-frequency slope of the compatibility defect is exactly the residual between ideal transfer slope and observed spectral-slope migration on a fixed frequency mask. Its pointwise depth curvature is also determined by a variance of log transfer gain under a depth-tilted spectral measure. The frozen benchmark contains 64 synthetic records and 65 nonoverlapping real-data blocks from electrocardiography, Bitcoin minute returns, and solar-wind magnetic-field data. Across 3483 retained cells, pooled R2 for ΔR4=R4−R3 ranges from 0.9891 to 0.9994 across the nine real dataset–estimator groups. After removing fixed-configuration means, R2 remains 0.6608–0.9943; across 81 fixed real configurations, the median R2 is 0.8782. The record includes the manuscript, frozen data, executed publication run, source code, dependency specification, provenance and licensing documentation, and SHA-256 manifests required to reproduce and audit the reported results. Reproducibility DOI: 10.5281/zenodo.22056027Corresponding author: gmtheory@outlook.fr Licensing is file- and source-specific. See DATA_LICENSES_AND_ATTRIBUTION.md for upstream licenses, attribution requirements, and provenance of the redistributed data.

Open access
2 source records
Solar and Space Plasma Dynamics
Cardiac Imaging and Diagnostics
Parallel Computing and Optimization Techniques
Original source
Aug 7, 2026·arXiv (Cornell University)
0 cites
Dual-Node NVIDIA DGX Spark over Tailscale: A Remote-Access Testbed for Distributed LLM Training and Cyber-Threat-Intelligence Fine-Tuning

Vasanth Iyer

Compact AI systems make local language-model experimentation increasingly accessible, yet practical evidence for multi-node training on desktop-class accelerators remains limited. This report presents a proof-of-concept deployment of distributed NanoChat pretraining across two NVIDIA DGX Spark systems, each with a GB10 Grace Blackwell system-on-chip and 128 GB of unified memory, administered remotely over a Tailscale mesh VPN and connected for training by a dedicated 200 Gb/s QSFP56 direct fiber link. PyTorch torchrun, DDP, and NCCL were configured with one process per node, a depth-20 NanoChat model, a local batch size of 32 per node, and a 2,048-token context, giving a global batch of 131,072 tokens per step. The run sustained a step time of about 69.4 s (about 1,890 tokens/s), processing about 653 million tokens over four days. We document link configuration, container setup, interface binding, a step-zero evaluation bug that triggered NCCL timeouts, checkpointing, and troubleshooting lessons, as a reproducibility reference for small labs. We also built a cybersecurity fine-tuning dataset from 77 CISA advisories (338 training, 37 validation conversations) and ran a 17-question held-out evaluation comparing a baseline SFT checkpoint against a CTI-augmented checkpoint with an Ollama-hosted LLM judge. CTI-specific categories improved while general-knowledge categories regressed, for a small overall change from 2.06 to 2.29 on a 0-10 scale. The same cluster supports a 400-level AI course (CS 426) and a query engine for CompTIA Security+ POGIL activities in CBS 255, showing modest local infrastructure can serve both research and teaching. The study establishes feasibility rather than a scaling-efficiency claim, since single-node throughput used for comparison was estimated, not measured under matched conditions. Runbook and scripts are available (see Code Availability).

Open access
Scientific Computing and Data Management
Parallel Computing and Optimization Techniques
Software System Performance and Reliability
Original source
Jun 12, 2026·IEEE Transactions on Parallel and Distributed Systems
0 cites
FZKP: Alleviating Dataflow Complexity to Exploit Fine-Grained Parallelism for ZKP Acceleration

Ziheng Xiao, Mingyu Yan, Mingyu Gao, Runzhen Xue · 7 authors

Zero-knowledge proof (ZKP) is a promising cryptographic protocol, but its practical deployment is hindered by the time-consuming proof generation. The proof generation inherently exhibits high-degree parallelism, yet challenges persist in exploiting fine-grained parallelism due to the dataflow complexity, impeding previous work to achieve optimal acceleration. In this work, we propose FZKP, a ZKP accelerator that utilizes two novel fine-grained dataflows coupled with two forward-flow microarchitectures to alleviate dataflow complexity, efficiently exploiting fine-grained parallelism. The proposed dataflows simplify the dataflow pattern for parallel execution, disclosing fine-grained parallelism at a low cost. The microarchitectures employ a base design to handle large bit-width intermediate results for timely consumption. They then replicate and combine the base design following the proposed dataflow to facilitate parallel execution. When evaluated in 12 nm, FZKP achieves an average speedup of 10.3× and 2.2× over the state-of-the-art GPU-based solution and ZKP accelerator on real-world workloads, respectively.

Parallel Computing and Optimization Techniques
Advanced Data Storage Technologies
Distributed systems and fault tolerance
Original source
May 25, 2026·arXiv (Cornell University)
0 cites
ZK-Tracer: A High-Performance Heterogeneous Accelerator for Zero-Knowledge VM Trace Generation

Jieran Cui, Zhengkai Wen, Haowen Fang, Yinan Zhu · 9 authors

Zero-knowledge virtual machines (zkVMs) are a key technology for driving the large-scale adoption of zero-knowledge proofs (ZKP), but their performance bottlenecks severely limit their practicality. While current hardware acceleration research has exclusively focused on backend proving, we identify that the frontend execution and trace generation phase is rapidly emerging as the new system bottleneck. To address this challenge, we propose ZK-Tracer, the first hardware accelerator architecture specifically designed for the zkVM frontend. ZK-Tracer features a novel heterogeneous design comprising a Main Trace Unit and parallel Permutation Trace Units. It exposes a fine-grained interface to the host software through a lightweight instruction set extension, enabling efficient task offloading. Our ASIC implementation results demonstrate that ZK-Tracer achieves up to 1829x speedup in trace generation over a high-performance multi-core CPU. When integrated with existing backend proving accelerators, it delivers a remarkable 963x end-to-end performance improvement for the entire ZKP system.

Open access
3 source records
cs.AR
Security and Verification in Computing
Cloud Computing and Resource Management
Original source
Apr 25, 2026·Zenodo (CERN European Organization for Nuclear Research)
0 cites
Accelerating ZK-Rollup Proof Generation 5.37× over Sequential Baselines: Modular Hypercube Chunking for L1-Resident Multi-Scalar Multiplication

Andrés Sebastián Pirolo

Abstract Multi-Scalar Multiplication (MSM) is the primary computational bottleneck in zero-knowledge (ZK) proof generation for decentralized networks. This research accelerates MSM by solving the memory bandwidth constraints inherent in high-dimensional elliptic curve cryptography. We introduce Modular Hypercube Chunking, a novel microarchitectural approach that partitions high-dimensional algebraic precomputations into smaller, orthogonal blocks. Specifically, we divide a 12-dimensional workload into three separate 4D hypercubes, restricting the entire memory footprint to 31.1 KB. This geometric partitioning ensures perfect residency within the ultra-fast L1 cache of modern processors. By employing shared doubling across these blocks, the algorithm processes twelve scalars simultaneously with a single elliptic curve duplication, bypassing slow RAM access entirely. Empirical evaluations conducted on an ARM Snapdragon 8 Gen 2 mobile processor demonstrate a peak 5.37× speedup compared to optimized sequential baselines, reducing the computational cost to 18.44 microseconds per scalar. These findings prove that geometric data partitioning within strict L1 cache boundaries significantly outperforms traditional arithmetic-heavy optimizations. The implications of this work provide a highly scalable architecture capable of executing server-grade ZK-Rollup proof generation on resource-constrained edge devices, while establishing a highly efficient blueprint for future multicore hardware accelerators. Furthermore, initial stress-tests of a 12D monolithic architecture (68 MB footprint) yielded an anomalous 8.88× peak speedup. This finding reveals a novel sparse-access memory optimization path, which we introduce as an open architectural challenge.

Open access
3 source records
Cryptography and Residue Arithmetic
Parallel Computing and Optimization Techniques
Polynomial and algebraic computation
Original source
Apr 20, 2026
0 cites
GPU Acceleration of the Sum-Check Protocol Over Towers of Binary Fields for Verifiable Computing

Andrew Fan, Yanze Wu, Harry Han, Md Tanvir Arafin

Emerging zero-knowledge proof protocols such as Binius and Binius-FRI operate over towers of binary fields, allowing for ultra-fast polynomial commitments over a base field. Sum-check, a key protocol in algebraic proof systems, is one of the key implementation bottlenecks for Binius and similar protocols. While sum-check is a massively parallel algorithm, GPU acceleration of sum-check has received little attention due to the lack of native GPU support for binary field multiplication. Hence, in this paper, we explore the key issues in existing GPU-based sum-check accelerators and present SumCATS - an efficient GPU implementation for sum-check acceleration. SumCATS leverages two fundamental improvements over the existing solutions. First, it adapts a CPU-based algorithmic improvement to sum-check proving and applies it to GPUs by recognizing the reduction pattern and shared memory optimizations. Secondly, SumCATS reduces the number of global memory accesses by precomputing products of random challenges and using base field operations to reconstruct extension field elements. When these optimizations are combined, SumCATS achieves a significant speedup (1.81× on NVIDIA RTX 3090 Ti, 1.62× on NVIDIA A100) over the baseline GPU implementation (Binius-GPU) for sum-check over binary tower fields. The code and research artifacts for SumCATS design are available at https://github.com/SPIRE-GMU/sum_cats.

Cryptography and Residue Arithmetic
Parallel Computing and Optimization Techniques
Security and Verification in Computing
Original source
Apr 7, 2026·Figshare
0 cites
OTIMIZAÇÃO DE GAS EM ETHEREUM: ANÁLISE DE OPCODES E ESTRUTURAS DE DADOS

Tiago Ferreira Cavazin

Este artigo analisa estratégias de otimização de gas em Ethereum a partir de duas dimensões principais: o custo dos opcodes da EVM e as escolhas de estruturas de dados em Solidity. A tabela de opcodes da EVM e a evolução do gas schedule mostram que operações de armazenamento e acesso externo, como SSTORE, SLOAD, CALL, BALANCE e EXT*, estão entre as mais caras, especialmente após EIPs como a 2929, que aumentaram o custo de acessos “frios” a contas e slots de storage para refletir melhor seu impacto na execução e na camada de armazenamento. Estudos recentes sobre custos de armazenamento evidenciam que uma escrita em SSTORE pode custar cerca de 22.100 gas para 32 bytes (aprox. 690 gas/byte), enquanto leituras via SLOAD também são significativamente caras, motivando pesquisas sobre técnicas como SSTORE2 e mecanismos para corrigir “overcharge” em leitura/escrita de storage, com ganhos médios de até 30–32% em fees para certos padrões de uso. Boas práticas de otimização de gas em Solidity incluem reduzir o número de acessos a storage movendo valores frequentemente lidos para variáveis em memória, empacotar variáveis em slots de 32 bytes (storage packing), preferir tipos fixos a dinâmicos quando possível, evitar cópias desnecessárias de arrays de storage para memória e desenhar estruturas de dados que minimizem gravações em storage. A literatura e guias de otimização indicam que a escolha entre arrays, mappings, structs e padrões de layout impacta diretamente o custo de execução, especialmente em loops que interagem com storage ou estruturas dinâmicas. Conclui‑se que a otimização de gas em Ethereum é um problema tanto de engenharia de baixo nível, ligado ao custo de opcodes e ao modelo warm/cold de acessos, quanto de design de dados e algoritmos, com implicações econômicas diretas para usuários, protocolos DeFi e estratégias de design de L2s.<br>

Open access
2 source records
Advanced Data Storage Technologies
Parallel Computing and Optimization Techniques
Optimization and Packing Problems
Original source
Apr 1, 2026·Blockchain Research and Applications
0 cites
PRISM: Provable and Immutable Storage Mechanism with Ethereum-based PDP

Shohei Kakei, Masanori Hirotomo, Masami Mohri, Yoshiaki Shiraishi

Highlights • Identifying threats that cannot be countered by theoretical security based on STRIDE threat analysis of an existing provable data possession (PDP) system • Designing a PDP system with practical security features to counter threats that cannot be addressed with theoretical security alone • Presenting the implementation of the proposed PDP system, PRISM, which is also provided as an open-source software • Validating security properties through property-based fuzz testing with 10,000 randomized test runs per security property • Demonstrating PRISM’s key strengths through comprehensive experiments, including basic performance, trade-offs between processing time and data auditing efficiency, and capabilities for detecting data anomalies Digital platforms are increasingly recognized as a cornerstone for advanced virtual spaces such as smart cities and the metaverse, where vast amounts of data are aggregated, analyzed, and utilized to make critical decisions. These platforms rely on data fusion to integrate diverse sources of information, encompassing individual behavior, urban dynamics, and system states. Through auditing against data tampering, loss, and substitution, enabling the detection of such threats is critical to building a highly reliable system. This paper introduces PRISM (Provable and Immutable Storage Mechanism), an Ethereum-based Provable Data Possession (PDP) system designed to integrate data reliability and security with decentralized auditing. PDP, a cryptographic protocol that enables data integrity in untrusted cloud storage, has seen extensive research focusing on theoretical security and computational efficiency. PRISM extends this foundation by addressing practical security concerns, including the integration of authentication and authorization, data immutability, data uniqueness, data freshness, and state management, to ensure a robust system implementation. Experiments on processing costs and parameter analysis reveal a trade-off between the costs and detection accuracy and demonstrate that PRISM provides efficient data auditing.

Open access
Advanced Data Storage Technologies
Distributed systems and fault tolerance
Parallel Computing and Optimization Techniques
Original source
Jan 1, 2026
0 cites
MEVisor: High-Throughput MEV Discovery in DEXs with GPU Parallelism

Weimin CHEN, Xiapu Luo

Decentralized finance (DeFi) is an emerging financial service on blockchain, enabling automatic and anonymous transactions.Within DeFi, decentralized exchanges (DEXs) maintain reserves of a pair of tokens and determine the exchange rate to swap tokens.However, DEXs also create opportunities for Maximal Extractable Value (MEV), where attackers include, exclude, or reorder DEX transactions to exploit price discrepancies of tokens and extract profit.Uncovering MEV opportunities requires high throughput, as the 12-second block interval and the vast search space impose strict time constraints.However, existing tools suffer from low throughput, as they rely on CPU-bound execution, which is hindered by frequent state forking and slow DEX execution.In this paper, we take the first step in leveraging GPU parallel computing power to boost MEV-search throughput in arbitrage and sandwich strategies.More precisely, we compile an MEV bot into a GPU application and then launch thousands of GPU threads to search for profit in parallel.To this end, we design new solutions to address three major challenges: designing cheatcodes to simulate transactions on GPU, proposing a memory manager to reduce GPU memory usage, and designing strategyaware mutations to improve input diversity.We implement a prototype named MeVisor that runs DEXs on GPUs and searches for MEV using a parallel genetic algorithm.Evaluated on 3,941 real MEV cases from Ethereum, MeVisor achieves 3.3M-5.1Mtransactions per second, outperforming the CPU baseline by 100,000x.In a large-scale study of Q1 2025 data, MeVisor estimates MEV opportunities ranging from 2 to 14 transactions, yielding at most $1.1 million in MEV profit.

Open access
Parallel Computing and Optimization Techniques
Embedded Systems Design Techniques
Advanced Neural Network Applications
Original source
Jan 1, 2026·IEEE Transactions on Information Forensics and Security
1 cites
Dishonest Majority Passive-to-Active Compiler Over Rings for MPC With Constant Online Communication

Jiandong Zhang, Han Jiang, Chenkai Zeng, Qi Feng · 8 authors

Secure multiparty computation (MPC) over Z2kis more efficient than computations over fields, and studying MPC protocols under malicious security has practical application value. Malicious security with a dishonest majority over rings remains challenging. The most popular approach is SPDZ2k, however, this is a specific protocol that does not support the transformation of any existing semi-honest MPC protocols into malicious security protocols. The zero knowledge proof (ZKP)-based compiler satisfies this requirement. Existing state-of-the-art protocols have logarithmic online communication overhead in terms of the circuit size |C|, and their direct application to rings is nontrivial as they were originally designed for finite fields. In this work, we investigate the communication overhead to develop malicious security protocols. We bridge the gap between malicious security with abort and semi-honest security, by constructing a “GMW-style” verification protocol to achieve malicious security in a dishonest majority setting. This approach incurs a constant online communication overhead by enhancing the machinery of zero-knowledge fully linear interactive oracle proof (zk-FLIOP). Additionally, we extend the zk-FLIOP to work over any ring by invoking reverse multiplication friendly embeddings (RMFEs). Our results show that the online communication complexity of the verification process depends on only the security parameter, the number of parties, and the ring size. Furthermore, for small-scale circuits over Z2, we designed a distributed lookup table argument where both the total communication complexity and the computational cost are independent of the circuit size but of the input wires.

Parallel Computing and Optimization Techniques
Distributed systems and fault tolerance
Numerical Methods and Algorithms
Original source
Jan 1, 2026·SSRN Electronic Journal
0 cites
Feasibility Study of Instruction-Level Pipelining within the Ethereum Virtual Machine Architecture

Gopal Ojha

The Ethereum Virtual Machine (EVM) is a stack-based virtual processor that executes smart contract bytecode sequentially. While this design ensures determinism and correctness, it inherently limits instruction throughput. This paper presents a feasibility study of instruction-level pipelining within the EVM interpreter architecture. By analyzing the internal execution flow of the EVM as implemented in the Go-Ethereum (geth) client, the study identifies the program counter dependency, particularly under jump instructions, as the principal control hazard preventing naïve pipelining. A two-stage pipelined execution model is proposed, separating opcode fetch and decode from execution and program counter update, with a feedback mechanism to preserve EVM semantics. The work focuses on architectural feasibility rather than performance evaluation and optimization, demonstrating that pipelining inside the EVM interpreter is conceptually possible under controlled synchronization. Limitations, design challenges, and future research directions are discussed.

Open access
2 source records
Security and Verification in Computing
Parallel Computing and Optimization Techniques
Cloud Computing and Resource Management
Original source
Dec 11, 2025·Productivity Press eBooks
0 cites
Tokenisation: The Basics

Kevin Wooldridge, Stephen Ashurst

Tokenisation is the process by which real-world assets or services are converted into digital tokens on a distributed ledger. The digital token becomes an on-chain representation of the real-world asset and can be managed by participants who have access to the distributed ledger as part of a blockchain network.

Developmental Biology and Gene Regulation
Genomics and Chromatin Dynamics
Parallel Computing and Optimization Techniques
Original source
Dec 7, 2025·Lecture notes in computer science
2 cites
Scalable zkSNARKs for Matrix Computations

Mingshu Cong, Sherman S. M. Chow, Siu Ming Yiu, Tsz Hon Yuen

No abstract is available for this record.

Numerical Methods and Algorithms
Matrix Theory and Algorithms
Parallel Computing and Optimization Techniques
Original source
Oct 18, 2025·ACM Transactions on Reconfigurable Technology and Systems
1 cites
HiFA: A High-Performance and Flexible Acceleration Framework for Large-Size Number Theoretic Transform

Qilin Hu, Haotian Wang, Chubo Liu, Keqin Li · 5 authors

Zero-Knowledge Proofs (ZKP) and Homomorphic Encryption (HE) are crucial for data privacy in applications like cloud, blockchain, and analytics. However, the real-world adoption often faces performance challenges, particularly in the execution of the Number Theoretic Transform (NTT) required for polynomial multiplication involving sizes beyond \(2^{20}\) and large integer widths (e.g., 256 bits). FPGAs offer a promising platform for acceleration, but efficiently implementing large-size NTTs remains difficult due to the limited on-chip resources. The widely adopted four-step NTT method, used to relieve the need for large on-chip memory, introduces performance bottlenecks. Initially, the traditional dataflow NTT architecture may not fully exploit available compute capability, which hinders achieving peak performance. Furthermore, during the matrix transpose phase, the non-sequential access to external High-Bandwidth Memory (HBM) causes inefficiency. To address these challenges, we introduce HiFA, an FPGA-based automatic accelerator framework designed for high-performance and flexible large-size NTT computations. HiFA utilizes a stacked NTT architecture for high parallelism, maximizing HBM throughput. It supports various decomposed polynomial sizes via a novel reordering module. Additionally, a specialized cyclic shuffle module is integrated to optimize data movement during the matrix transpose step, alleviating random memory access delay. HiFA also provides an automatic Design Space Exploration (DSE) framework that identifies optimal four-step decomposition parameters and generates corresponding hardware configurations. Our experiments show that the FPGA implementation of HiFA achieves an average speedup of 2.97× and up to 7.25× improvement in latency over prior state-of-the-art FPGA solutions. Compared to prior GPU-based methods, HiFA achieves an average energy efficiency gain of 2.24×.

Open access
Algorithms and Data Compression
Chaos-based Image/Signal Encryption
Parallel Computing and Optimization Techniques
Original source
Sep 17, 2025·arXiv (Cornell University)
2 cites
ZKProphet: Understanding Performance of Zero-Knowledge Proofs on GPUs

Tarunesh Verma, Yichao Yuan, Nishil Talati, Todd Austin

Zero-Knowledge Proofs (ZKP) are protocols which construct cryptographic proofs to demonstrate knowledge of a secret input in a computation without revealing any information about the secret. ZKPs enable novel applications in private and verifiable computing such as anonymized cryptocurrencies and blockchain scaling and have seen adoption in several real-world systems. Prior work has accelerated ZKPs on GPUs by leveraging the inherent parallelism in core computation kernels like Multi-Scalar Multiplication (MSM). However, we find that a systematic characterization of execution bottlenecks in ZKPs, as well as their scalability on modern GPU architectures, is missing in the literature. This paper presents ZKProphet, a comprehensive performance study of Zero-Knowledge Proofs on GPUs. Following massive speedups of MSM, we find that ZKPs are bottlenecked by kernels like Number-Theoretic Transform (NTT), as they account for up to 90% of the proof generation latency on GPUs when paired with optimized MSM implementations. Available NTT implementations under-utilize GPU compute resources and often do not employ architectural features like asynchronous compute and memory operations. We observe that the arithmetic operations underlying ZKPs execute exclusively on the GPU's 32-bit integer pipeline and exhibit limited instruction-level parallelism due to data dependencies. Their performance is thus limited by the available integer compute units. While one way to scale the performance of ZKPs is adding more compute units, we discuss how runtime parameter tuning for optimizations like precomputed inputs and alternative data representations can extract additional speedup. With this work, we provide the ZKP community a roadmap to scale performance on GPUs and construct definitive GPU-accelerated ZKPs for their application requirements and available hardware resources.

Open access
3 source records
Cryptography and Residue Arithmetic
Cryptography and Data Security
Polynomial and algebraic computation
Original source
Aug 22, 2025·arXiv
1 cites
zkPHIRE: A Programmable Accelerator for ZKPs over HIgh-degRee, Expressive Gates

Alhad Daftardar, Jianqiao Mo, Joey Ah-kiow, Benedikt Bünz · 6 authors

Zero-Knowledge Proofs (ZKPs) have emerged as a powerful tool for secure and privacy-preserving computation. ZKPs enable one party to convince another of a statement's validity without revealing anything else. This capability has profound implications in many domains, including machine learning, blockchain, image authentication, and electronic voting. Despite their potential, ZKPs have seen limited deployment because of their exceptionally high computational overhead, which manifests primarily during proof generation. To mitigate these overheads, a (growing) body of researchers has proposed hardware accelerators and GPU implementations of both kernels and complete protocols. Prior art spans a wide variety of ZKP schemes that vary significantly in computational overhead, proof size, verifier cost, protocol setup, and trust. The latest and widely used ZKP protocols are intentionally designed to balance these trade-offs. One particular challenge in modern ZKP systems is supporting complex, high-degree gates using the SumCheck protocol. We address this challenge with a novel programmable accelerator to efficiently handle arbitrary custom gates via SumCheck. Our accelerator achieves upwards of $1000\times$ geomean speedup over CPU-based SumChecks across a range of gate types. We include this unit in zkPHIRE, a programmable, full-system accelerator that accelerates the HyperPlonk protocol. zkPHIRE achieves $1486\times$ geomean speedup over CPU and $11.87\times$ geomean speedup over the state-of-the-art at iso-area. Together, these results demonstrate compelling performance while scaling to large problem sizes (upwards of $2^{30}$ constraints) and maintaining small proof sizes ($4-5$ KB).

Open access
2 source records
cs.AR
cs.CR
Parallel Computing and Optimization Techniques
Original source
Jun 22, 2025
0 cites
ALLMod: Exploring Area-Efficiency of LUT-based Large Number Modular Reduction via Hybrid Workloads

Fangxin Liu, Haoming Li, Zongwu Wang, Bo Zhang · 8 authors

Modular arithmetic, particularly modular reduction, is widely used in cryptographic applications such as homomorphic encryption (HE) and zero-knowledge proofs (ZKP). High-bit-width operations are crucial for enhancing security; however, they are computationally intensive due to the large number of modular operations required. The lookup-table-based (LUT-based) approach, a “space-for-time” technique, reduces computational load by segmenting the input number into smaller bit groups, pre-computing modular reduction results for each segment, and storing these results in LUTs. While effective, this method incurs significant hardware overhead due to extensive LUT usage. In this paper, we introduce ALLMod, a novel approach that improves the area efficiency of LUT-based largenumber modular reduction by employing hybrid workloads. Inspired by the iterative method, ALLMod splits the bit groups into two distinct workloads, achieving lower area costs without compromising throughput. We first develop a template to facilitate workload splitting and ensure balanced distribution. Then, we conduct design space exploration to evaluate the optimal timing for fusing workload results, enabling us to identify the most efficient design under specific constraints. Extensive evaluations show that ALLMod achieves up to $\lt sup\gt1\lt/sup\gt|.65 \times$ and $3 \times$ improvements in area efficiency over conventional LUT-based methods for bit-widths of 128 and 8,192, respectively.

Advanced Data Storage Technologies
Distributed and Parallel Computing Systems
Parallel Computing and Optimization Techniques
Original source
May 12, 2025·2025 IEEE Symposium on Security and Privacy (SP)
3 cites
ZHE: Efficient Zero-Knowledge Proofs for HE Evaluations

Zhelei Zhou, Yun Li, Yuchen Wang, Zhaomin Yang · 8 authors

Homomorphic Encryption (HE) allows computations on encrypted data without decryption. It can be used where the users' information are to be processed by an untrustful server, and has been a popular choice in privacy-preserving applications. However, in order to obtain meaningful results, we have to assume an honest-but-curious server, i.e., it will faithfully follow what was asked to do. If the server is malicious, there is no guarantee that the computed result is correct. The notion of verifiable HE (vHE) is introduced to detect malicious server's behaviors, but current vHE schemes are either more than four orders of magnitude slower than the underlying HE operations (Atapoor et. al, CIC 2024) or fast but incompatible with server-side private inputs (Chatel et. al, CCS 2024). In this work, we propose a vHE framework ZHE: efficient Zero-Knowledge Proofs (ZKPs) that prove the correct execution of HE evaluations while protecting the server's private inputs. More precisely, we first design two new highly-efficient ZKPs for modulo operations and (Inverse) Number Theoretic Transforms (NTTs), two of the basic operations of HE evaluations. Then we build a customized ZKP for HE evaluations, which is scalable, enjoys a fast prover time and has a non-interactive online phase. Our ZKP is applicable to all Ring-LWE based HE schemes, such as BGV and CKKS. Finally, we implement our protocols for both BGV and CKKS and conduct extensive experiments on various HE workloads. Compared to the state-of-the-art works, both of our prover time and verifier time are improved; especially, our prover cost is only roughly 27–36× more expensive than the underlying HE operations, this is two to three orders of magnitude cheaper than state-of-the-arts.

2 source records
Parallel Computing and Optimization Techniques
Numerical Methods and Algorithms
Original source
May 10, 2025·INTERANTIONAL JOURNAL OF SCIENTIFIC RESEARCH IN ENGINEERING AND MANAGEMENT
0 cites
Web Based Hierarchical Deterministic wallet

Naval Kishor Jha

Abstract Pixel-Web3 Wallet is a hierarchical deterministic (HD) wallet designed for secure and decentralized asset management across multiple blockchain networks, including Ethereum and Solana. Unlike traditional wallets that depend on browser extensions or centralized servers, Pixel offers a web-based solution with user-controlled security through locally stored seed phrases. This paper explores the wallet’s architecture, security framework, and innovative features, such as real-time balance updates and flexible recovery options. Additionally, the research evaluates the scalability of Pixel and its potential expansion to support more blockchain networks. By eliminating reliance on third-party services, Pixel enhances accessibility while maintaining strong security, making it a promising solution for blockchain enthusiasts, traders, and developers. Keywords: Blockchain, HD Wallet, Cryptocurrency,Web3,Ethereum,Solana, Security

Open access
Parallel Computing and Optimization Techniques
Mobile Agent-Based Network Management
Distributed and Parallel Computing Systems
Original source
Apr 9, 2025·arXiv (Cornell University)
0 cites
Conthereum: Concurrent Ethereum Optimized Transaction Scheduling for Multi-Core Execution

Atefeh Zareh Chahoki, Maurice Herlihy, Marco Roveri

Conthereum is a concurrent Ethereum solution for intra-block parallel transaction execution, enabling validators to utilize multi-core infrastructure and transform the sequential execution model of Ethereum into a parallel one. This shift significantly increases throughput and transactions per second (TPS), while ensuring conflict-free execution in both proposer and attestor modes and preserving execution order consistency in the attestor. At the heart of Conthereum is a novel, lightweight, high-performance scheduler inspired by the Flexible Job Shop Scheduling Problem (FJSS). We propose a custom greedy heuristic algorithm, along with its efficient implementation, that solves this formulation effectively and decisively outperforms existing scheduling methods in finding suboptimal solutions that satisfy the constraints, achieve minimal makespan, and maximize speedup in parallel execution. Additionally, Conthereum includes an offline phase that equips its real-time scheduler with a conflict analysis repository obtained through static analysis of smart contracts, identifying potentially conflicting functions using a pessimistic approach. Building on this novel scheduler and extensive conflict data, Conthereum outperforms existing concurrent intra-block solutions. Empirical evaluations show near-linear throughput gains with increasing computational power on standard 8-core machines. Although scalability deviates from linear with higher core counts and increased transaction conflicts, Conthereum still significantly improves upon the current sequential execution model and outperforms existing concurrent solutions under a wide range of conditions.

Open access
2 source records
cs.CR
cs.DC
Distributed and Parallel Computing Systems
Original source