• QRAKEN achieves strict F1 scores of 0.643 with GPT-4.1 mini and 0.652 with GPT-5.4 on the CK25 benchmark, delivering relative gains of 30% and 32% over the strongest recomputed participant.
arXiv NLP · Knowledge Graphs
Foresight-over-Graph: Reasoning Beyond Local Horizons for Knowledge Base Question Answering
• Foresight-over-Graph achieves state-of-the-art performance on knowledge base question answering benchmarks, improving Hit on the CWQ dataset by 16.58% while reducing language model calls and token usage.
arXiv NLP · Knowledge Graphs
MAWARITH: A Dataset and Benchmark for Legal Inheritance Reasoning with LLMs
• MAWARITH is a new dataset containing 12,500 Arabic legal inheritance cases, designed to train and evaluate LLMs on complex, multi-step reasoning for Islamic inheritance law. • The dataset supports the full reasoning chain, including identifying heirs, applying rules, and calculating exact shares, offering step-by-step solutions and justifications.
arXiv NLP · Knowledge Graphs
Language as a Wave Phenomenon: Semantic Phase Locking and Interference in Neural Networks
• The PRISM model, a complex-valued encoder, demonstrates that semantic relationships correlate with phase structure, with synonym pairs showing significantly higher phase coherence (R=0.198) than random pairs (R=0.072).
arXiv NLP · Knowledge Graphs
• The Geometric Mixture-of-Experts (GeoMoE) framework fuses node representations across diverse Riemannian spaces using Ollivier-Ricci Curvature (ORC) to model complex graph topologies. • GeoMoE employs a graph-aware gating network with node-specific weights, regularized by a curvature-guided alignment loss for interpretable and geometry-consistent routing.
arXiv AI · Knowledge Graphs
JFTA-Bench: Evaluate LLM's Ability of Tracking and Analyzing Malfunctions Using Fault Trees
• The JFTA-Bench benchmark, containing 3130 entries and averaging 40.75 turns per entry, evaluates LLMs on malfunction localization using fault trees in multi-turn dialogues. • A novel textual representation of fault trees is proposed to enable direct processing by LLMs, aiding in malfunction tracking and analysis.
arXiv AI · Knowledge Graphs
6 stories
BMFM-RNA: whole-cell expression decoding improves transcriptomic foundation models
• Whole-cell expression decoding (WCED) outperforms masked language modeling (MLM) in transcriptomic foundation models for downstream tasks, despite higher training reconstruction error. • WCED reconstructs the entire gene vocabulary from a single CLS token embedding, enabling improved cell representations by creating a maximally informative bottleneck.
arXiv AI · Knowledge Graphs
• A multi-view selector increases direct pair recall from 53.74% to 76.98% at 72,660 candidate edges, yet cluster recall only rises marginally from 32.54% to 33.33%. • The evaluation on a GLEIF sample of 3,633 names and 2,880 source identities shows that 202 out of 208 newly retrieved silver-positive pairs fall below the decision threshold.
arXiv AI · Knowledge Graphs
Aligning Multimodal Patient Evidence with Biomedical Knowledge Graphs for Clinical LLMs
• The MM-KG framework explicitly links multimodal patient observations with biomedical knowledge graphs using route-prioritized alignment, yielding drug-controlled AUROC interaction improvements of +0.194 on MIMIC-IV and +0.299 on ADNI datasets.
arXiv NLP · Knowledge Graphs
Wikidata Search Traces: A Dataset for Training Knowledge Graph Search Agents
• Language model agents improve knowledge graph search accuracy by utilizing a recursive language model harness that batches graph calls and manages retrieved evidence in persistent Python state rather than direct tool calling. • The researchers construct multi-hop questions on a frozen Wikidata snapshot and release 10,235 solving traces alongside the harness to train graph search agents.
arXiv NLP · Knowledge Graphs
Curriculum Brain: Constructing Curriculum Knowledge Graphs as a Substrate for Cognitive Diagnosis
• Curriculum Brain automates the construction of curriculum knowledge graphs by pairing a version-controlled knowledge base with an agentic pipeline of eleven single-responsibility agents. • The system processes official curriculum documents and textbooks to generate and validate concept-skill mappings, achieving a 91.5% resolution rate without human escalation in later pipeline configurations.
arXiv AI · Knowledge Graphs
EgoSelf: From Memory to Personalized Egocentric Assistant
• EgoSelf introduces a graph-based interaction memory that captures temporal and semantic relationships from past observations to construct user-specific profiles for personalized egocentric assistants.
arXiv AI · Knowledge Graphs
Why Better Cross-Lingual Alignment Fails for Better Cross-Lingual Transfer: Case of Encoders
• Explicit cross-lingual alignment techniques fail to improve token-level downstream performance because alignment and downstream task objectives are largely orthogonal. • Analysis of four XLM-R encoder models shows embedding distances are unreliable predictors of task performance improvements, and alignment and task gradients often conflict.
arXiv NLP · Knowledge Graphs
Meta-Reinforcement Learning with Self-Reflection for Agentic Search
• Meta-Reinforcement Learning with Self-Reflection (MR-Search) trains agents to adapt search strategies across episodes by generating explicit self-reflections. • This approach improves in-context exploration at test-time by leveraging self-reflections as additional context for subsequent search attempts.
arXiv NLP · Knowledge Graphs
8 stories
Enrich-on-Graph: Query-Graph Alignment for Complex Reasoning with LLM Enriching
• Enrich-on-Graph (EoG) framework achieves state-of-the-art performance on knowledge graph question answering benchmarks by leveraging large language models to enrich knowledge graphs and bridge the semantic gap with unstructured queries.
arXiv NLP · Knowledge Graphs
Continual Graph Memory for Mathematical Research Agents
• Ansatz successfully closes all ten tasks in the First Proof Second Batch benchmark by utilizing a Continual Graph Memory system to organize long-horizon mathematical proof searches. • The system employs a unified graph architecture alongside an evidence-sensitive curator to manage intermediate facts, plans, and counterexamples across parallel agent explorations.
arXiv AI · Knowledge Graphs
• The Hop-Decayed Influence attack achieves an 88 to 94 percent success rate across HotpotQA and 2WikiMultiHopQA benchmarks while corrupting only 0.016 percent of auxiliary structures in Microsoft GraphRAG and HippoRAG2 architectures.
arXiv AI · Knowledge Graphs
• Asterism extracts observations from hundreds of papers as concept-relation triples and unifies those concepts within a hierarchical ontology. • A field deployment with ten researchers demonstrates that users curate an evidence graph and aggregate observations at various levels of granularity to form theories aligned with their preferences.
arXiv NLP · Knowledge Graphs
Improving Scientific Document Retrieval with Academic Concept Index
• An academic concept index, organized by taxonomy, is introduced to address limitations in general-domain retriever adaptation for scientific documents. • The academic concept index enhances query generation via CCQGen for broader concept coverage and context augmentation with CCExpand for concept-focused snippets.
arXiv AI · Knowledge Graphs
ComplianceNLP: Knowledge-Graph-Augmented RAG for Multi-Framework Regulatory Gap Detection
• ComplianceNLP, a new system integrating a knowledge-graph-augmented RAG pipeline and multi-task obligation extraction, achieves 87.7 F1 on regulatory gap detection, outperforming GPT-4o+RAG by 3.5 F1. • The system processes 9,847 regulatory updates over four months, demonstrating 96.0% estimated recall and 90.7% precision while increasing analyst efficiency by 3.1x.
arXiv NLP · Knowledge Graphs
DeFAb: A Verifiable Benchmark for Defeasible Abduction in Foundation Models
• The DeFAb benchmark, converting knowledge bases into formally grounded instances for defeasible abduction, generates over 372,648 instances from 18 sources, with rule-based solvers achieving 100% accuracy while frontier language models reach a maximum of 65%.
arXiv AI · Knowledge Graphs
• Relational semantics emerge in autoregressive LLMs with sufficient logic-bearing supervision, even in shallow 2-3 layer models, as demonstrated by a controlled Knowledge Graph-based synthetic framework. • Successful generalization to unseen entities in relational tasks aligns with stable intermediate-layer signals within the LLMs.
arXiv NLP · Knowledge Graphs
8 stories
TRN-R1-Zero: Text-rich Network Reasoning via LLMs with Reinforcement Learning Only
• TRN-R1-Zero is a post-training framework for text-rich network (TRN) reasoning that is trained solely via reinforcement learning, directly optimizing base LLMs with a Neighbour-aware Group Relative Policy Optimisation objective.
arXiv NLP · Knowledge Graphs
M$^3$KG-RAG: Multi-hop Multimodal Knowledge Graph-enhanced Retrieval-Augmented Generation
• M$^3$KG-RAG enhances multimodal retrieval-augmented generation by constructing a multi-hop multimodal knowledge graph (M$^3$KG) and implementing GRASP for precise entity grounding and selective pruning.
arXiv NLP · Knowledge Graphs
Let Them Steal: Trapping Large Language Model Extraction Attacks with Knowledge Honeypot
• Knowledge Trap, a new defense for large language models, redirects extraction attacks towards low-transferability knowledge using a Honeypot Knowledge Graph and breadcrumb-guided exploration. • Experiments in medical and financial domains demonstrate that Knowledge Trap reduces surrogate agreement by an average of 6.2% without degrading legitimate-user accuracy.
arXiv AI · Knowledge Graphs
• XMedFusion, a modular AI framework, enhances autonomous medical systems by decomposing visual information into functional components including a visual perception agent, a knowledge graph construction agent, and a retrieval-guided drafting process.
arXiv AI · Knowledge Graphs
4 stories
Knowledge-Based Zero-Replay Debugging of Multi-Agent LLM Traces
• A new knowledge-based approach frames multi-agent LLM trace debugging as a decision-support problem, compiling traces into structured event knowledge graphs. • The BranchPoint-Latent predictor, calibrated against a replay oracle on 37 trace families, improves per-trace localization (Branch Recall@5) from 0.73 to 0.93 at zero replay cost.
arXiv AI · Knowledge Graphs
Knowledge-Guided Manipulation Using Multi-Task Reinforcement Learning
• The Knowledge Graph based Massively Multi-task Model-based Policy Optimization (KG-M3PO) framework unifies perception, knowledge, and policy for robotic manipulation in partially observable settings by augmenting vision with an online 3D scene graph.
arXiv AI · Knowledge Graphs
GLiNER-Relex: A Unified Framework for Joint Named Entity Recognition and Relation Extraction
• GLiNER-Relex introduces a unified architecture that extends the GLiNER framework to perform both named entity recognition (NER) and relation extraction (RE) within a single model. • The model leverages a shared bidirectional transformer encoder and enables zero-shot extraction of arbitrary entity and relation types specified at inference time.
arXiv NLP · Knowledge Graphs
• Frontier LLMs exhibit a 'Provenance Gap,' fabricating citations, with the best model achieving only 15.3% relevant PubMed identifiers even when prompted. • The HEG-TKG system, built from 4,512 PubMed records and curated sources, achieves 100% evidence verifiability with 203 inline citations for rare disease reasoning, matching baseline clinical feature coverage.
arXiv NLP · Knowledge Graphs
4 stories
LLM-Assisted Discovery of Typed Semantic Links for Ontology Network Construction
• An automated framework for ontology network construction combines domain-adapted DistilBERT embeddings, clustering-based pre-filtering, and GPT-4o-driven relationship generation to reduce 800,000 raw concept pairs to 95,000 high-quality candidates across 33 ontologies in ReproduceMeON.
arXiv NLP · Knowledge Graphs
Ontology-Grounded, Reasoner-Verified Benchmarks for Evaluating LLM Reasoning in Scientific AI
• A novel automated pipeline generates ontology-grounded multiple-choice question benchmarks from OWL 2 ontologies to evaluate logical reasoning in scientific AI applications. • The method produces formally verified distractors by perturbing class definition axioms and using an OWL reasoner to check entailment across datasets including 112 Pizza, 2491 PMDco, and 15216 DOID questions.
arXiv AI · Knowledge Graphs
• CoDHy introduces an interactive AI co-scientist system that generates biomarker-guided drug combination hypotheses for oncology applications. • The system builds task-specific knowledge graphs from curated databases and biomedical literature, integrating graph embeddings with agent-based reasoning to validate and rank drug combinations.
arXiv NLP · Knowledge Graphs
Build2SPARQL: A Large-Scale Text-to-SPARQL Benchmark Dataset for Building Knowledge Graph Querying
• Build2SPARQL provides 6,136 executable SPARQL queries and 30,680 natural-language questions derived from 201 building knowledge graphs using ontologies like Brick and ASHRAE 223P.
arXiv AI · Knowledge Graphs
Protecting De-identified Documents from Search-based Linkage Attacks
• A novel method protects de-identified documents from search-based linkage attacks by identifying and rewriting N-grams present in fewer than k documents. • The approach employs an inverted index for efficient N-gram analysis followed by an LLM-based rewriter to reformulate sensitive spans.
arXiv NLP · Knowledge Graphs
ESIA: An Energy-Based Spatiotemporal Interaction-Aware Framework for Pedestrian Intention Prediction
• The ESIA (Energy-based Spatiotemporal Interaction-Aware framework) proposes a novel Conditional Random Field (CRF)-based paradigm for pedestrian intention prediction, treating pedestrians and the environment as spatiotemporal nodes in a unified graph.
arXiv AI · Knowledge Graphs
• The Multi-Agent Knowledge Analysis (MAKA) architecture separates intent routing, quantitative analysis, knowledge retrieval, and verification to support risk-aware human-AI decision-making in manufacturing.
arXiv AI · Knowledge Graphs
TusoAI: Agentic Optimization for Scientific Methods
• TusoAI, an agentic AI system, autonomously develops and optimizes computational methods for scientific tasks by integrating domain knowledge into a knowledge tree and performing iterative, domain-specific optimization.
arXiv AI · Knowledge Graphs
8 stories
GraphCert: Bootstrap Agentic Graph Reasoning with Certified Evidence Rubrics
• GraphCert outperforms significantly larger LLM agents and existing post-training methods across five distinct graph reasoning domains in GRBENCH. • The method utilizes a Bootstrapped Graph Quizzer guided by generation controls to produce graph-grounded QA pairs, which undergo execution certification and semantic curation to form certified evidence rubrics.
arXiv AI · Knowledge Graphs
• Researchers introduce the Concept Lifecycle Model and Concept-Grounded Attention to integrate graph-grounded, temporally versioned entities with explicit provenance into transformer computations. • Evaluations on MuSiQue and HotpotQA demonstrate that explicit temporal representation improves answer accuracy by 13 to 25 points across generators up to 122B parameters.
arXiv AI · Knowledge Graphs
• A new Semantic Web-based framework consolidates fragmented evidence for Fundamental Rights Impact Assessments under Article 27 of the EU AI Act into a SPARQL-queryable knowledge graph of 1,351 RDF triples.
arXiv AI · Knowledge Graphs
• The paper proposes Riemannian Foundation Models (RFMs) as a new paradigm for graph intelligence, leveraging Riemannian geometry to overcome limitations of current Graph Neural Networks (GNNs) and Large Language Models (LLMs) in capturing complex graph structures.
arXiv AI · Knowledge Graphs
• Chain of Retrieval (COR) is a new iterative framework for full-paper retrieval that decomposes query papers into aspect-specific views and iteratively expands search results. • COR aggregates retrieval results in a post-order manner, recursively merging descendant query nodes with their parent nodes to capture hierarchical relations.
arXiv NLP · Knowledge Graphs
Rule-State Inference (RSI): A Bayesian Framework for Compliance Monitoring in Rule-Governed Domains
• Rule-State Inference (RSI) is a Bayesian framework that models compliance monitoring by inferring latent rule activation, compliance rates, and parametric drift from partial observations. • RSI achieves a 600x speedup in adapting to regulatory changes compared to full model retraining, absorbing updates in under 1ms.
arXiv AI · Knowledge Graphs
Machine Translation in the Wild: User Reaction to Xiaohongshu's Built-In Translation Feature
• Xiaohongshu's built-in translation feature, launched in January 2025, receives generally positive user reactions, with notable appreciation for translating posts and comments. • User testing of the translation function extends beyond standard text to include diverse inputs like internet slang, abbreviations, emoji, kaomoji, and coded texts.
arXiv NLP · Knowledge Graphs
7 stories
• GeoOutageBench integrates visual, textual, and structured data from power outage records, remote sensing, weather observations, and domain ontologies to evaluate large language model performance on geospatiotemporal knowledge graph question answering.
arXiv AI · Knowledge Graphs
Neural Structural Reasoner: A Brain-inspired Architecture for Reasoning over Structured Knowledge
• Neural Structural Reasoner preserves relational structure directly in the connectivity and dynamics of coupled neuronal populations to perform link prediction on knowledge graphs. • The architecture implements biological mechanisms including multi-layered encoding, stable representations, and path integration to parallelize computation over candidate relational structures.
arXiv AI · Knowledge Graphs
ImbalancE: Inference-Time Latent Search Against Degree Imbalance in Link Prediction
• Link prediction errors in Knowledge Graph Embedding models strongly correlate with the degree imbalance between anchor and target entities in test triples. • The proposed ImbalancE method utilizes an inference-time latent search optimization technique to explore the embedding space and blend out-of-band information at evaluation time.
arXiv AI · Knowledge Graphs
Triples and Knowledge-Infused Embeddings for Clustering and Classification of Scientific Documents
• Abstract-only inputs achieve the strongest and most stable classification performance on a filtered arXiv corpus, reaching 0.923 accuracy and 0.923 macro-F1 across a five-seed benchmark (seeds 40-44).
arXiv NLP · Knowledge Graphs
• A dynamic framework using multimodal data analytics, a hierarchical knowledge graph with adaptive edge weighting, and heterogeneous graph attention combined with temporal sequence modeling predicts learning behavior and identifies at-risk students in advanced mathematics.
arXiv AI · Knowledge Graphs
• SceneDiver, a method for vision-language decision making, generates focus plans by first building a scene graph for comprehension and then iteratively decomposing tasks through recognition, understanding, and analysis.
arXiv AI · Knowledge Graphs
An LLM-Based System for Argument Reconstruction
• A novel LLM-based system reconstructs arguments from natural language text into abstract argument graphs with premises, conclusions, and support/attack/undercut relations. • The system's performance is evaluated through manual analysis on a textbook dataset and quantitative comparison on benchmark datasets, demonstrating adequate recovery of argumentative structure.
arXiv NLP · Knowledge Graphs
7 stories
APOLO: Automatic Prompt Optimization for Ontology Learning
• APOLO introduces an automatic prompt optimization framework for ontology learning using large language models, validated on the biomedical DOID and plant PO ontologies. • The methodology utilizes a multi-agent system to generate text-ontology training pairs and employs GEPA, a greedy evolutionary prompt optimizer built on DSPy, to train greedy and autoregressive learner architectures.
arXiv AI · Knowledge Graphs
5W1H+Which: Context-Valid Semantic Indexing with Progressive Ontology Binding
• Researchers introduce 5W1H+Which, a semantic indexing framework that decouples source-grounded content extraction from versioned ontology binding to preserve query flexibility and support formal reasoning.
arXiv NLP · Knowledge Graphs
Scalable GNN-based Knowledge Graph Representation Learning with Efficient Message Passing
• Researchers introduce an extended Relational Sparse Matrix Multiplication framework that supports expressive composition functions such as 2x2 block-diagonal matrix multiplication, Givens rotation, and circular correlation for knowledge graph representation learning.
arXiv AI · Knowledge Graphs
LLM-Guided Ontology-Driven Knowledge Graph Construction from Unstructured Text
• Schema-guided prompting significantly improves entity and relation extraction quality when building ontology-driven knowledge graphs from a private corpus of 80 power-grid incident reports.
arXiv NLP · Knowledge Graphs
• The Eolas pipeline uses large language models to automatically transform scientific documents into knowledge graphs aligned with a specified ontology, significantly reducing extraction time from 30-90 minutes to a few minutes.
arXiv AI · Knowledge Graphs
Knowledge Graph-Enhanced Zero-Shot Topic Classification: A Multi-Strategy Comparative Study
• A zero-shot multi-label topic classification framework augmented with per-article knowledge graphs is proposed and evaluated across fifteen LLMs and eight multi-label datasets. • Keyword-enhanced classification emerges as the top-performing base method, with six LLMs outperforming a sentence-encoder baseline without graph augmentation.
arXiv NLP · Knowledge Graphs
LLM-Assisted Ontology Engineering and Construction of a French Legal Knowledge Graph
• A two-stage LLM-assisted workflow is presented for constructing a French legal knowledge graph from maintenance regulations using GPT-4.1 and mistral-large-2512. • The methodology involves open extraction of entities and triples, label normalization via embedding-based fusion, and induction of candidate object properties, followed by closed extraction guided by the resulting ontology.
arXiv AI · Knowledge Graphs
Initial Evaluation of Potential Bias in Coverage of Humans in Wikidata
• Women account for 28.71% (CI +/-0.04) of all humans in Wikidata with a stated gender. • Western Europe and North America represent approximately 53% of citizenship statements among humans in Wikidata. • Rural birthplaces are observed in 2.48% of classified Wikidata entries, significantly lower than the 27.4% global baseline.
arXiv NLP · Knowledge Graphs
8 stories
FlyAOC: Evaluating Agentic Ontology Curation of Drosophila Scientific Knowledge Bases
• FlyAOC evaluates AI agents on end-to-end ontology curation using a benchmark of 7,397 expert-curated annotations across 100 genes from FlyBase. • Agents utilize a corpus of 16,898 papers along with ontology resources to search for evidence and recover structured annotations such as function terms and expression patterns.
arXiv NLP · Knowledge Graphs
Do LLMs Understand Context? A Knowledge Graph-Based Evaluation Framework
• A novel knowledge graph-based evaluation framework called Semantic Structural Similarity for KGs (S3KG) achieves F1 gains of up to plus 7.6 points and an AUROC up to 0.973 across nine benchmarks. • The framework combines structural and semantic signals into a single score and incorporates a diagnostic analysis system to categorize reasoning errors at the triplet level.
arXiv AI · Knowledge Graphs
HCOE: Hyperbolic Clinical Ontology Embeddings from Biomedical Language Models
• Hyperbolic Clinical Ontology Embeddings map frozen BioBERT representations into a Poincare ball to explicitly preserve medical code hierarchies for biomedical language models. • The approach combines parent-side and child-side ontology-guided contrastive learning with coarse-to-fine ontology-path aggregation using ICD, CCS, and ATC classification systems.
arXiv AI · Knowledge Graphs
MedMASLab: A Unified Orchestration Framework for Benchmarking Multimodal Medical Multi-Agent Systems
• MedMASLab provides a unified framework and benchmarking platform for multimodal medical multi-agent systems, integrating 11 heterogeneous architectures across 24 medical modalities. • The framework includes a standardized communication protocol, an automated clinical reasoning evaluator using LLMs for zero-shot semantic evaluation, and the most extensive benchmark to date covering 473 diseases.
arXiv AI · AI & Machine Learning
• WeatherTGD employs a training-free, multi-agent framework that leverages Text Gradient Descent (TGD) for interpretable weather time series captioning. • The system uses three specialized LLM agents (Statistical Analyst, Physics Interpreter, Meteorology Expert) whose textual gradients are fused by a Consensus-Aware Gradient Fusion mechanism.
arXiv NLP · AI & Machine Learning
Gender Dynamics and Homophily in a Social Network of LLM Agents
• A social media platform composed entirely of autonomous AI chatbots, with over 70,000 agents and 140 million posts over one year, reveals fluid gender performance among agents. • Despite gender fluidity, agents exhibit strong homophily, consistently following others performing similar genders, driven by both social selection and social influence mechanisms.
arXiv AI · AI & Machine Learning
Do LLMs Understand Collaborative Signals? Diagnosis and Repair
• Large Language Models (LLMs) can outperform traditional matrix factorization models in recommendation tasks when provided with collaborative user-item interaction data in a clear, easily digestible format.
arXiv NLP · AI & Machine Learning
7 stories