This paper asks what must be added to three-dimensional semantic segmentation before a distributed biological system can be said to organize differentiated, object-specific content. It begins with ordinary object perception and separates class labels, instance identity, border ownership, recurrent completion, multisensory registration, receiver state, Phase Wave Differentials, action, and returned sensory correction. The paper introduces a fifty-equation formal specification and an Object-Boundary Registration Benchmark. A deterministic reference application, fitted synthetic pilots, distribution-shift tests, latent and global alternatives, targeted ablations, calibration analysis, six evidence figures, and a machine-checked finite contract kernel make the proposal auditable. The synthetic results are mixed: typed receiver structure outperforms a global summary, while stronger latent alternatives match or exceed it under some noise and missingness conditions. Those adverse results remain central to the paper. The work therefore presents a testable research program, not completed biological or consciousness validation. A staged biological protocol is frozen, but it has not been run and the final test remains sealed. The public companion archive contains the complete cumulative manuscripts, source and claim ledgers, executable application, tests, structured results, negative-result record, figure provenance, formal proof receipts, and reproducibility instructions.
Open access
2 source records
Cell Image Analysis Techniques
Face Recognition and Perception
Generative Adversarial Networks and Image Synthesis
Computational research depends on the ability to independently reproduce results, yet modern workflows are fragile: they drift across environments, depend on undocumented assumptions, and often fail silently. ValiChord provides a decentralised, agentâcentric infrastructure for independent reproducibility validation. Validators reâexecute workflows in diverse environments, generate cryptographically signed attestations, and contribute structured detector evidence that captures environment drift, dependency skew, execution variability, and workflow fragility. A commitâreveal protocol preserves validator independence, while Harmony Records synthesise divergent outcomes without collapsing them into binary judgements. ValiChord validates computation, not data provenance, and is explicit about this boundary: it strengthens the computational layer of scientific integrity without claiming to detect data fabrication. The system is built on Holochain, not blockchain, ensuring tamperâevident provenance without global ledgers, tokens, or consensus mechanisms. Reference implementation and detector suite: https://github.com/topeuph-ai/ValiChord
Abstract The peer-reviewed journal article imposes structural constraints on the dissemination, validation, and reuse of research outputs. Intermediate results, negative findings, methodological refinements, and replication attempts are systematically underrepresented in published literature, limiting visibility into ongoing research activity for both scientists and mission-driven funders. Here we present Carrierwave, an open infrastructure for continuous, granular scientific communication built on structured research objects (ROs), cryptographic provenance, blockchain-based attribution, and programmable incentive mechanisms. Each RO represents an atomic unit of scientific output -- a single experimental result, negative finding, dataset, protocol, or replication -- that is hashed for content integrity, stored in a persistent database, and optionally minted as an ERC-721 non-fungible token on the Ethereum blockchain. The system includes an on-chain bounty pool enabling funders to directly incentivize specific research activities, and an automated analysis layer that synthesizes disclosed ROs into continuously updated research landscape maps. We describe the system architecture, report on its implementation and deployment on Ethereum mainnet, and present a quantitative analysis of disease-specific publication frequency demonstrating the information latency problem that Carrierwave addresses. The distribution of publication frequency across disease areas is highly skewed, with the majority of conditions represented by fewer than four publications per year in high-impact biology journals. For diseases in the long tail, the interval between successive publications may span months or years. Publication frequency correlates poorly with disease burden, instead reflecting historical research community size and advocacy momentum. By reducing the unit of communication to the individual research object and eliminating editorial gatekeeping as a prerequisite for disclosure, Carrierwave increases the effective sampling rate of scientific activity in precisely the domains where publication-based visibility is most sparse. The system is live at https://carrierwave.org .
Single-cell foundation models (scFMs)-transformer networks pretrained by self-supervision on tens of millions of single-cell transcriptomes-have moved rapidly from proof of concept to a central methodological theme in computational biology. Yet much of the literature evaluates them on the same downstream tasks (cell-type annotation, batch integration, perturbation prediction) where strong, inexpensive classical baselines already exist, and on several of these tasks the foundation-model advantage is modest or contested. This review takes a different framing: rather than asking whether scFMs win every benchmark, we ask what they offer that task-specific and classical methods structurally cannot. We identify and analyze six comparative advantages: (i) label-efficient transfer and zero-/few-shot inference from a single pretrained backbone; (ii) atlas-scale generalization and reference-free integration across datasets, tissues, and technologies; (iii) a unified multi-task, multi-omic interface that amortizes engineering and modeling effort; (iv) context-dependent, attention-derived gene and cell embeddings that enable network inference and in silico perturbation; (v) predictable scaling behavior with data, parameters, and compute; and (vi) cross-species and cross-modality knowledge transfer, including the interplay between what protein language models already encode and what genuinely requires single-cell pretraining. For each advantage we summarize the supporting evidence, the limits exposed by recent benchmarks and linear-baseline critiques, and the open questions. We conclude that the durable value proposition of scFMs is reusability and breadth-a single artifact that transfers across problems-rather than uniform state-of-the-art accuracy, and we outline what would strengthen the case for that proposition.
BACKGROUND: An image sharing framework is important to support downstream data analysis especially for pandemics like Coronavirus Disease 2019 (COVID-19). Current centralized image sharing frameworks become dysfunctional if any part of the framework fails. Existing decentralized image sharing frameworks do not store the images on the blockchain, thus the data themselves are not highly available, immutable, and provable. Meanwhile, storing images on the blockchain provides availability/immutability/provenance to the images, yet produces challenges such as large-image handling, high viewing latency while viewing images, and software inconsistency while storing/loading images. OBJECTIVE: This study aims to store chest x-ray images using a blockchain-based framework to handle large images, improve viewing latency, and enhance software consistency. BASIC PROCEDURES: We developed a splitting and merging function to handle large images, a feature that allows previewing an image earlier to improve viewing latency, and a smart contract to enhance software consistency. We used 920 publicly available images to evaluate the storing and loading methods through time measurements. MAIN FINDINGS: The blockchain network successfully shares large images up to 18 MB and supports smart contracts to provide code immutability, availability, and provenance. Applying the preview feature successfully shared images 93% faster than sharing images without the preview feature. PRINCIPAL CONCLUSIONS: The findings of this study can guide future studies to generalize our framework to other forms of data to improve sharing and interoperability.
Approaches developed based on the blockchain concept can provides a framework for the realization of open science. The traditional centralized way of data collection and curation is a labor-intensive work that is often not updated. The fundamental contribution of developing blockchain format of microbial databases includes: 1. Scavenging the sparse data from different strain database; 2. Tracing a specific thread of access for the purpose of evaluation or even the forensic; 3. Mapping the microbial species diversity; 4. Enrichment of the taxonomic database with the biotechnological applications of the strains and 5. Data sharing with the transparent way of precedent recognition. The plausible applications of constructing microbial databases using blockchain technology is proposed in this paper. Nevertheless, the current challenges and constraints in the development of microbial databases using the blockchain module are discussed in this paper.
In this issue of Cytometry A, Zhao et al. (page 1073â1080) report on their work to diagnose leukemic B cell non-Hodgkin's Lymphoma from flow cytometry (FCM) raw data of blood and bone marrow samples using a dedicated computer approach, which would assign one of eight B-cell lymphoma diagnoses or ânormalâ to a sample. A remarkable of level of classification performance could be achieved in the validation set. For the âtrueâ classification of B-cell lymphomas, conventional diagnostics had incorporated morphology, FCM and additional information from histology and genetics if needed, whereas computer diagnosis was derived from FCM data alone. In this context, uncertainty to delineate, for example, monoclonal B-cell lymphocytosis from chronic lymphocytic leukemia or to subclassify a B-cell malignancy as either mantle cell lymphoma or prolymphocytic leukemia is not an outright error, but is rather based on the limitations of FCM itself. Furthermore, cell populations tagged as abnormal by the algorithm and color-coded accordingly in conventional plots can help human diagnosticians to review and fine-tune the diagnosis. However, some lymphomas (most prominent in follicular lymphoma) were classified as normal by the algorithm. Vice versa, only few samples classified as ânormalâ by human diagnosticians were classified as lymphoma by the algorithm. Thus, a deficit in sensitivity exists, which is clinically relevant. Computer support is instrumental for the analysis of FCM data, because nobody is able to draw conclusions from raw list mode files. However, conventional FCM computer programs execute relative simple tasks to support the workflow of a human researcher or diagnostician. In a typical workflow, several sequential steps have to be performed (Fig. 1, left side). Fluorescence spillover compensation is calculated from control samples. One-dimensional transformation of raw data (logarithmic, logical, possibly a shift of zero and negative values to some defined minimum, etc.) is routinely performed on fluorescence channels. Data are displayed in histograms or two-dimensional plots. Starting gates are used to look for artifacts and to remove debris and cells not of interest. A considerable number of plots are necessary, if several fluorochromes are used and several populations are of interest. Data from several samples with identical panel may be displayed in parallel in an overlay. Cells are tagged according to gates in these plots and may then be displayed separately and/or color-coded. Hierarchical and/or Boolean gating strategies are used for the definition of cell populations and subpopulations of interest. Cell numbers and antigen expression of these cell populations of interest constitute the readout of a single tube. A final result or diagnosis is derived assessing this readout or the synopsis of the readout of several tubes. All of the calculations in such a manual workflow are based on straight âif A then Bâ logic, performing calculations on a maximum of two parameters concurrently. Conventional FCM computer support aims at displaying data in a clear manner to the human operator, especially effects of manipulation in two-parameter plots upon plots of other parameters, but not at automation. The most advanced process in standard applications is the calculation of fluorescence spillover compensation, which nowadays usually is performed in some (semi-) automated fashion. However, although every single step in this procedure is quite straightforward, due to the multitude of plots and gates from current 10 to 14 parameter FCM data, important information may be missed. In the recent decades, many attempts have been reported to introduce more advanced computation methods into histology, cytopathology, image cytometry and conventional FCM analysis (1, 2). These algorithms will be called artificial intelligence (AI) from here, although some of them do not deserve this name in its strict sense. Two strategies, sometimes overlapping, are applied in these attempts: firstly, AI may be used to automatize conventional data processing and analysis as described above in order to reduce the workload for the investigator, reduce bias using standardized procedures, and speed up analyses. To this end, regarding FCM, algorithms search for minimal values in distributions to define optimal positions for gates to divide populations or search for appropriate cut off values to gate out debris. Furthermore, normalization algorithms can be applied to level out differences due to instrument settings or biological variations in sets of multiple similar data. Many of these algorithms are available in the Bioconductor âflow Coreâ FCM package implemented in R (3). Secondly, new methods were introduced that go beyond the sequential analysis of two-dimensional plots and base calculations on more parameters of the higher-dimensional space in parallel, which is a crucial need nowadays, when standard cytometers report 10 to 14 parameters per cell and dedicated research instruments up to over 100 parameters. Such algorithms can either substitute conventional strategies, for example, to gate cell populations and read out antigen expression levels or they can be used to extract information from the raw data that is not accessible by conventional gating (4). One of the prominent tasks within an FCM workflow is to define cell populations within a mixture of different cells (âclusteringâ) that may be of interest for research or diagnosis. AI can directly use higher dimensional data as input for cell clustering or it can perform dimensionality reduction and data visualization, for example, by tSNE or one of its variants (5, 6) or SOM (7), the latter already including some clustering of the data. After dimensionality reduction, population clustering can be added by separate AI algorithms or a human operator can take over for this task, integrating the output of the dimensionality reduction and conventional gating. Many different algorithms are able to solve the task of clustering in an automated fashion either performing a two-step procedure integrating dimension reduction and subsequential clustering or direct clustering of higher dimensional data; however, as shown in the FlowCAP challenges, results are not unequivocal, especially, if the number of clusters is not defined a priori, and differences remain between different algorithms and human experts. Up to now, no perfect automatic solution for cell clustering exists, although many solutions perform quite well (7). Furthermore, clustering revealing further information on relatedness between populations has been suggested for a multitude of different research questions, for example, cellular developmental trajectories, and has been optimized according to these special tasks (further Ref. in 4). Furthermore, metadata extracted from raw FCM data may also be clustered, for example, in order to define diagnostic or prognostic subgroups (8). Whereas unsupervised clustering can be helpful for many exploratory research questions to identify cell populations and subpopulations, for medical diagnostic purposes supervised AI methods have been described, that use external information such as diagnoses or outcome to train the AI, for example, using support vector machines or neural networks. All of these strategies rely on a large dataset for training and may incorporate more or less steps from a conventional workflow (4, 9, 10). Manual gating and tagging of cell populations may be used for training of the AI (11) or AI may be trained using only the final results, that is, diagnosis, as described, for example, in Ref. (12) or in the work by Zhao et al. discussed here. Several AI strategies have been able to discern overt acute myeloid leukemia (AML) from normal samples with a high success rate in the second FlowCap challenge (7), however, this can be a considered a quite simple task, since overt AML is easily characterized by a large abnormal population of blast or sometimes monocytic cells. In contrast, separation of AML from myelodysplastic syndromes or from acute lymphoblastic leukemia, everyday questions in diagnostics, is less trivial. In contrast to simplified âyes or noâ tasks, Zhao et al. tackled a much more realistic question: to deduce a specific diagnosis from FCM panels as they are used in conventional diagnostics. They achieved this goal without an attempt to mimic a conventional human FCM workflow. They transformed the FCM data by self-organizing maps (SOM) and classified these representations by a convolutional neural network (CNN), dealing with each tube separately first and finally with data from all three tubes. The researchers took advantage of a very large database of patient sample FCM data. Data from more than 18,000 samples analyzed in a uniform fashion with identical antibody combinations and more than 200 samples of the rarest subtype of lymphoma could be used to train the CNN. In order to get some insight into the CNN âblack box,â they checked, which markers were of most importance for the AI to classify a specific diagnosis correctly and they had cell populations tagged that were detected to be abnormal and discriminative by the algorithm for the respective disease in a way to understand the AI's decision (and to use this assignment for a possible refinement by a human diagnostician in practical diagnostic use in the future). As described above, the results of their approach are remarkable, but a problem in sensitivity to detect all true lymphoma cases remains, which is most prominent for follicular lymphoma. Maybe the CNN could be trained in a way, that the correct distinction B-NHL of any type versus normal is assigned a higher weight compared to B-NHL subtyping. If we inspect the importance of single markers for AI performance in Supporting Figure 5, we note that some diagnosis assignments rely heavily on a few markers, whereas other diagnoses seem to rather depend on the distribution of many markers. Interestingly, the latter diagnoses without dependence on dominant markers have the highest rate of falsely being categorized as normal (follicular lymphoma, marginal zone lymphoma, lymphoplasmactic lymphoma). Furthermore, for a human diagnostician, an imbalance of kappa versus lambda light chain expression on B cells is a very important clue for a diagnosis of B-cell lymphoma, whereas the CNN of Zhao et al. does not seem to rely heavily on this information. In a different approach, to detect minimal residual disease in childhood acute leukemia, conventional gating was used to train a machine learning algorithm based on Gaussian mixture models (11). Thus, for the non-AI expert the idea comes up, if some information of a conventional workflow, collected by an automated application, could be âinjectedâ into a CNN algorithm. If we assume that the problem of sensitivity will be tackled by improved versions in the near future, the AI solution of Zhao et al. will in fact be able to perform at âhematologist-levelâ and may even deliver B-NHL subtyping competence exceeding the results of conventional FCM alone. However, further problems have to be solved for a broader uptake of such a method: different laboratories work with different antibody panels and even antibodies recognizing the same cluster of differentiation antigen behave differently due to different antibody clones, different fluorochromes and different spillover from other fluorochromes in the panel. Thus, some methods of knowledge transfer are needed, if we want to avoid starting again with a training sample of more than 10,000 cases for every new antibody panel. If researchers will be able to solve these problems, AI for diagnostic FCM may finally leave the âproof of conceptâ stage and enter routine diagnostics. Open access funding enabled and organized by Projekt DEAL.
CĂŒneyt GĂŒrcan Akçora, Yitao Li, Yulia R. Gel, Murat KantarcıoÄlu
Proliferation of cryptocurrencies (e.g., Bitcoin) that allow pseudo-anonymous transactions, has made it easier for ransomware developers to demand ransom by encrypting sensitive user data. The recently revealed strikes of ransomware attacks have already resulted in significant economic losses and societal harm across different sectors, ranging from local governments to health care. Most modern ransomware use Bitcoin for payments. However, although Bitcoin transactions are permanently recorded and publicly available, current approaches for detecting ransomware depend only on a couple of heuristics and/or tedious information gathering steps (e.g., running ransomware to collect ransomware related Bitcoin addresses). To our knowledge, none of the previous approaches have employed advanced data analytics techniques to automatically detect ransomware related transactions and malicious Bitcoin addresses. By capitalizing on the recent advances in topological data analysis, we propose an efficient and tractable data analytics framework to automatically detect new malicious addresses in a ransomware family, given only a limited records of previous transactions. Furthermore, our proposed techniques exhibit high utility to detect the emergence of new ransomware families, that is, ransomware with no previous records of transactions. Using the existing known ransomware data sets, we show that our proposed methodology provides significant improvements in precision and recall for ransomware transaction detection, compared to existing heuristic based approaches, and can be utilized to automate ransomware detection.
Davi R. Ortega, Catherine M. Oikonomou, H. Ding, Prudence Rees-Lee · 6 authors
Abstract Three-dimensional electron microscopy techniques like electron tomography provide valuable insights into cellular structures, and present significant challenges for data storage and dissemination. Here we explored a novel method to publicly release more than 11,000 such datasets, more than 30 TB in total, collected by our group. Our method, based on a peer-to-peer file sharing network built around a blockchain ledger, offers a distributed solution to data storage. In addition, we offer a user-friendly browser-based interface, https://etdb.caltech.edu , for anyone interested to explore and download our data. We discuss the relative advantages and disadvantages of this system and provide tools for other groups to mine our data and/or use the same approach to share their own imaging datasets.
Open access
2 source records
Innovative Microfluidic and Catalytic Techniques Innovation