Обработка текстовой информации включает редактирование, форматирование, проверку орфографии и лингвистический анализ текстов.


487 публикаций

Нажмите рядом со статьёй — скопируете ссылку для списка литературы по ГОСТ.

Your Mouse and Eyes Secretly Leak Your Preference: LLM Alignment using Implicit Feedback from Users
Beyond Global Replanning: Hierarchical Recovery for Cross-Device Agent Systems
StylisticBias: A Few Human Visual Cues Drive Most Social Biases in MLLMs
LedgerAgent: Structured State for Policy-Adherent Tool-Calling Agents
Rethinking Reward Supervision: Rubric-Conditioned Self-Distillation
Learning User Simulators with Turing Rewards
Native Active Perception as Reasoning for Omni-Modal Understanding
On the Memorization Behavior of LLMs in Generative Recommendation: Observations, Implications, and Training Strategies
Designing Recommendation Exposure and Favorite Lists: A Field Experiment in a Spot-Work Platform
RSRank: Learning Relevance from Representational Shifts
Temporal Preference Optimization for Unsupervised Retrieval
Decoupling Search from Reasoning: A Vendor-Agnostic Grounding Architecture for LLM Agents
Querit-Reranker: Training Compact Multilingual Rerankers via Efficient Label-Free Distribution Adaptation
Which Sections of a Research Paper Best Reveal Its Research Methods? Evidence from Library and Information Science
Output Vector Editing for Memorization Mitigation in Large Language Models
SAMA: Semantic Anchor-aligned Augmentation for Unified Low-Resource Multimodal Information Extraction
Lost in a Single Vector: Improving Long-Document Retrieval with Chunk Evidence Aggregation
RedactionBench
GraphPO: Graph-based Policy Optimization for Reasoning Models
Mitigating Scoring Errors and Compensating for Nonverbal Subtests in Speech-Based Dementia Assessment
Beyond Tokenization: Direct Timestep Embedding and Contrastive Alignment for Time-Series Question Answering
G-IdiomAlign: A Gloss-Pivoted Benchmark for Cross-Lingual Idiom Alignment
Evaluating Prompting-Based Defenses Against Domain-Camouflaged Injection Attacks
CEO-Bench: Can Agents Play the Long Game?
Fair Cognitive Impairment Detection Through Unlearning
Speech-Driven End-to-End Language Discrimination towards Chinese Dialects
Retrievable Gradients: Continual Post-Training Without Cumulative Weight Drift
One Sequential Recommendation Model Pretrained from Synthetic Priors Predicts Multiple Datasets
Intelligent Multimodal Retrieval and Reasoning for Geospatial Knowledge Discovery on the I-GUIDE Platform
MAGE-RAG: Multigranular Adaptive Graph Evidence for Agentic Multimodal RAG in Long-Document QA
Entity Labels Are Not Entity Signals: A Framework for Observable Relevance in Document Re-Ranking
Theorem-Grounded Execution Ontologies for Interpretable Machine Reasoning
Viral Images: Identifying Reprintings within 1.5 Million Photographs in Chronicling America
RL-Index: Reinforcement Learning for Retrieval Index Reasoning
Visuals Lie, Consistency Speaks: Disentangling Spatial Attention from Reliability in Vision-Language Models
NarrativeWorldBench: A Frontier-Saturated Benchmark and a Latent World Model for Long-Horizon Co-Creative Audio Drama
Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
Retrieval-as-a-Service:A System-Oriented Analysis of Industrial Retrieval Pipelines in Web Systems
VoidPadding: Let [VOID] Handle Padding in Masked Diffusion Language Models so that [EOS] Can Focus on Semantic Termination
Reading between the Lines: Leveraging Large Language Models for Global Dementia and Depression Assessment from Clinical Interviews
Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients
Darshana Graph: A Parallel Commentary Corpus for Comparative Indian Philosophy, with Stylometric and Exploratory Graph Analyses
ReproRepo: Scaling Reproducibility Audits with GitHub Repository Issues
Variable-Width Transformers
Scaling LLM Reasoning from Minimal Labels: A Semi-Supervised Framework with a Lightweight Verifier
Understanding the Behaviors of Environment-aware Information Retrieval
Tying the Loop -- Tied Expert Layers in Mixture-of-Experts Language Models
A Critical Evaluation of the Application of Natural Language Processing (NLP) For Enhancing Text and Speech Mining
KVEraser: Learning to Steer KV Cache for Efficient Localized Context Erasing
Benchmarking LLM Agents on Meta-Analysis Articles from Nature Portfolio
Context-Aware RL for Agentic and Multimodal LLMs
The Value Axis: Language Models Encode Whether They're on the Right Track
Can LLM Agents Infer World Models? Evidence from Agentic Automata Learning
Uncertainty Is Not a Safety Net for Clinical VQA, but Can It Anticipate Model Failure?
SING: Synthetic Intention Graph for Scalable Active Tool Discovery in LLM Agents
How Far Can Machine Translation Quality Take You? Extrinsic Discourse Evaluation in Goal-Oriented Setups
How Fine-Grained Should a RAG Benchmark Be? A Hierarchical Framework for Synthetic Question Generation
Semantic Identification of IoT Devices from Behavioral Primitives
Personalization and Evaluation of Conversational Information Access
ChronoID: Infusing Explicit Temporal Signals into Semantic IDs for Generative Recommendation
ScoreGate: Adaptive Chunk Selection for Retrieval-Augmented Generation via Dual-Score Statistical Fusion
Verifiable User Simulation for Search and Recommendation Systems
Private Information Retrieval for Large-Scale DNA-Based Data Storage
Non-Parametric Machine Text Detection via Multi-View Gaussian Processes
Harsher on Male? Evaluating LLMs on Gender-Asymmetric Moral Framing Across Diverse Conflict Scenarios
Diffusion-Refined Segmentation and Vision-Language Interpretation for Pediatric Brain Tumor MRI
Simulating Students' Java Programming Errors with Large Language Models
Poker Arena: Multi-Axis Profiling of Strategic Reasoning and Memory in LLMs
When Plausible Is Not Realistic: Evaluating Human Mobility in LLM-Based Urban Simulation
Hybrid Classical-Quantum Variational Autoencoder for Neural Topic Modeling
SuperThoughts: Reasoning Tokens in Superposition
Achieving Precise Text-To-Cypher Via Grounded Knowledge Graph Data Generation
Detecting Historical Turning Points in Italian Media: A Complex Systems Approach to a Diachronic News Corpus
Be My Tutor: On-Policy Co-Distillation for Mutual LLM Improvement via Peer Feedback
Learning to Hear Hesitation: Continual Learning for Disfluency-Aware ASR
AdaSR: Adaptive Streaming Reasoning with Hierarchical Relative Policy Optimization
Persona-Pruner: Sculpting Lightweight Models for Role-Playing
ClinHallu: A Benchmark for Diagnosing Stage-Wise Hallucinations in Medical MLLM Reasoning
Gaze Heads: How VLMs Look at What They Describe
DeRes: Decoupling Residual Stability and Adaptivity for Scalable CTR Prediction
GIScholarBench: Benchmarking LLM Overconfidence in GIS Research
Have I Solved This Before? Retrieving Similar Segmentation Problems for Evolutionary Learning
EmpiriGraph-Psy: A Dataset and LLM Pipeline for Extracting Empirical Relation Graphs from Psychology Abstracts
DeMix: Debugging Training Data with Mixed Data Error Types by Investigating Influence Vectors
The Long Tail, Not the Front Page: Cold-Start Prediction of Crowd Highlight Salience
CompRank: Efficient LLM Reranking via Token-Level Compression and Decoding-Free Scoring
FAST-MEL: A Fast, Accurate, and Storage Efficient Solution for Multimodal Entity Linking
MaxProof: Scaling Mathematical Proof with Generative-Verifier RL and Population-Level Test-Time Scaling
SupraBench: A Benchmark for Supramolecular Chemistry
Leveraging Audio-LLMs to Filter Speech-to-Speech Training Data
When Does Mixing Help? Analyzing Query Embedding Interpolation in Multilingual Dense Retrieval
HyperTool: Beyond Step-Wise Tool Calls for Tool-Augmented Agents
Influcoder: Distilling Decoders' Gradient Influence Rankings into an Encoder for Data Attribution
Learning to Reason by Analogy via Retrieval-Augmented Reinforcement Fine-Tuning
EvoArena: Tracking Memory Evolution for Robust LLM Agents in Dynamic Environments
A Context-Aware Dataset for Stance Detection in Bioethical Controversies on Reddit
SICI: A Semantic-Pragmatic Complexity Index Reveals Regime Shifts in LLM Stance Detection
Stability in Competitive Search with Results Diversification
Mult-DPO: Multinomial Direct Preference Optimization for Recommender Systems
MetaPlate: Counterfactual-Guided RAG-LLM Tool for Personalized Food Recommendation and Hyperglycemia Prevention
$τ$-Rec: A Verifiable Benchmark for Agentic Recommender Systems
CoDeR: Local Constraint-Compatible Retrieval Beyond Semantic Similarity
TimeLens: On-Device Artifact Recognition with Retrieval-Augmented Question Answering for the Grand Egyptian Museum
CQC-RAG: Robust Retrieval-Augmented Generation via Cross-Query Consistency
OneRetrieval: Unifying Multi-Branch E-commerce Retrieval with an Editable Generative Model
MDForge: Agentic Molecular Dynamics Pipeline Design under Sparse Simulator Feedback
Polar: A Benchmark for Evaluating Political Bias in LLMs
Order Is Not Control
Multi-Turn Reasoning When Context Arrives in Pieces: Scalable Sharding and Memory-Augmented RL
Natural Language Processing and Text Mining
Sentiment Analysis of the Top 5 E-commerce Platforms in Indonesia using Text Mining and Natural Language Processing (NLP)
FLOWREADER: Min-Cost Flow Optimization for Multi-Modal Long Document Q&A
Constrained Dominant Sets for Multimodal Document Question Answering
Gated Bidirectional Linear Attention for Generative Retrieval
Report on CHIIR 2026 Workshop on Generative AI and Academic Search (GAI&AS)
EviProp: Seeded Relevance Diffusion on Chunk-Page Graphs for Long Multimodal Document Retrieval
Dummy Backdoor as a Defense: Removing Unknown Backdoors via Shared Internal Mechanisms for Generative LLMs
The Long Tail, Not the Front Page: Cold-Start Prediction of Crowd Highlight Salience
Can AI Reason Like an Urban Planner? Benchmarking Large Language Models Against Professional Judgment
Organize then Retrieve: Hierarchical Memory Navigation for Efficient Agents
DiffCold: A Diffusion-based Generative Model for Cold-Start Item Recommendation
Findings of the MAGMaR 2026 Shared Task
A Controlled Study of Decoding-Time Truthfulness Methods on Instruction-Tuned LLMs
OpenMedReason: Scientific Reasoning Supervision for Medical Vision-Language Models
A Resource for Enthymeme Detection in Controversial Political Discourse
Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application
System Report for CCL25-Eval Task 5: New Dataset and LoRA-Fine-Tuned Qwen2.5
Redesign Mixture-of-Experts Routers with Manifold Power Iteration
Doc-to-Atom: Learning to Compile and Compose Memory Atoms
Context-Driven Incremental Compression for Multi-Turn Dialogue Generation
Fine-tuning Multi-modal LLMs with ART: Art-based Reinforcement Training
I Understand How You Feel: Enhancing Deeper Emotional Support Through Multilingual Emotional Validation in Dialogue System
Beyond representational alignment with brain-guided language models for robust reasoning
Notes2Skills: From Lab Notebooks to Certainty-Aware Scientific Agent Skills
Beyond Patches: Superpixel Token-based Transformers for Attribute-Specific Fashion Retrieval
Effective Reinforcement Learning for Agentic Search by Recycling Zero-Variance Queries During Training
miniReranker: Efficient Multimodal Reranking through Visual Cache Reuse and Interaction Sparsity
ConvMemory v2: A Recall-Preserving Top-10 Evidence Reranker for Conversational Memory Retrieval
Personal Salience: Highlighting Is Social, but Individuality Lives in Selection
Decoy-Calibrated Failure Audits for Language Models
Dynamic Linear Attention
Speaker Group Encoding in Self-supervised Speech Recognition Models
Are We Evaluating Knowledge or Phrasing? Mitigating MCQA Sensitivity with ParaEval
Multilingual Word-Level Forced Alignment with Self-Supervised Representations and Learned Dynamic Programming
Generative Archetype-Grounded Item Representations for Sequential Recommendation
Parallel Causal Associative Fields: Gated Sparse Memory for Long-Context Language Modeling
Enhancing Multilingual LLM-based ASR with Mixture of Experts and Dynamic Downsampling
SpenseGPT: Practical One-shot Pruning Enabling Sparse and Dense GEMMs for LLM Inference
Teach Multimodal Recommendation Model to See via Personalized Visual Extraction and Adaptive Learning
Provenance-Grounded Gating and Adaptive Recovery in Synthetic Post-Training Data Curation
Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models
iOSWorld: A Benchmark for Personally Intelligent Phone Agents
Data Synthesis and Parameter-Efficient Fine-Tuning for Low-Resource NMT: A Case Study on Q'eqchi' Mayan
SIGA: Self-Evolving Coding-Agent Adapters for Scientific Simulation
Causally Evaluating the Learnability of Formal Language Tasks
Driving Video Retrieval for Complex Queries with Structured Grounding
Closing the Indexing-Decoding Gap in Multimodal Generative Retrieval via Prefix Retention Optimization
Popcorn: A Configurable Benchmark for Visual Evidence in Multimodal Movie Recommendation
Toward Signing Activity Projection in Sign Language Interaction
Guide Me Out: A Framework to Benchmark VLM Operators Communication in Crisis Scenarios
MUDIDI: A Two-Stage Framework for Multilingual Dictionary Digitization with Language Models
Reasoning without Gold Standards: A Proxy-Judge Theory of Autoformalization
Sentiment Analysis of Social Media Using Natural Language Processing (NLP)
A Vision-language Framework for Comparative Reasoning in Radiology
Mind the Gap: Bridging Behavioral Silos with LLMs in Multi-Vertical Recommendations
Symbolic and Abstractive Reasoning with Complex Visual Queries
The Injection Paradox: Brand-Level Suppression in Safety-Trained LLM Recommendations via RAG Context Injection
TruthSplit: Operationalizing Conditional Validity in Arguments Through Multi-Perspective Reasoning
One Model, Multiple Goals: Adaptive Multi-Objective Learning for E-commerce Dialogue Systems
TABVERSE: Benchmarking Cross-Format Table Understanding in LLMs and VLMs
Clinically Grounded Privacy Evaluation of Medical LMs
Automated IEP Generation from Traditional Chinese Parent-Teacher Interviews via Corpus-Grounded Feature Diffusion
AGENTSERVESIM: A Hardware-aware Simulator for Multi-Turn LLM Agent Serving
Emergent Misalignment Can Be Induced by Sycophancy and Reversed via Alignment Gating
A Unifying Lens on Reward Uncertainty in RLHF
Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy
MAAM: Anchor-Preserving Compression and Contextual Calibration for Chinese Discriminatory Language Detection
Gated Bidirectional Linear Attention for Generative Retrieval
PaperFlow: Profiling, Recommending, and Adapting Across Daily Paper Streams
Tree-of-Experience: A Structured Experience-Management Solution for Self-Evolving Agents under Low-Repetition and Implicit-Reward Environments
Contrastive Training with LLM-generated Near-Misses for Robust Code-Switching Speech Recognition
Principles of Concept Representation in Sentence Encoders
RASFT: Rollout-Adaptive Supervised Fine-Tuning for Reasoning
Cartridges at Scale: Training Modular KV Caches over Large Document Collections
Distributional Approximate Nearest Neighbour Search for Uncertainty-Aware Retrieval
QO-Bench: Diagnosing Query-Operator-Preserving Retrieval over Typed Event Tuples
Improving the Efficiency and Effectiveness of LLM Knowledge Distillation for Conversational Search
Bradley-Terry Rankings for Recommender Systems Across Dataset Taxonomies
Your UnEmbedding Matrix is Secretly a Feature Lens for Text Embeddings
Agentopia: Long-Term Life Simulation and Learning in Agent Societies
PromptPrint: Behavioral Biometrics Through Natural Language Prompting in LLMs
A Four-Condition Diagnostic Protocol for Evidence Utilization in Long-Context and Retrieval-Augmented Language Models
When Better Codebooks Are Not Enough: Predictive Performance and Behavioral Reliability in LLM Political Event Coding
Explain Like I'm 5 or Whatever I Choose: Evaluating the Interactive Potential of Language Model Responses
From Correctness to Utility: Gain-Based Prefix Evaluation for LLM Reasoning
HKVM-RAG: Key-Value-Separated Hypergraph Evidence Organization for Multi-Hop RAG
Adversarial Creation and Detection of AI-Generated Social Bot Content
DEFINED: A Data-Efficient Computational Framework for Fine-Grained Creativity Assessment in Debate Scenarios
EviRank: Evidence-Based Confidence Estimation for LLM-Based Ranking
Archi: Agentic Operations at the CMS Experiment
From Self to Other: Evaluating Demographic Perspective-Taking in LLM Hate Speech Annotation
Many Circuits, One Mechanism: Input Variation and Evaluation Granularity in Circuit Discovery
BEATS: Bootstrapping E-commerce Attribute Taxonomies for Search through Iterative Human-AI Collaboration
MLEvolve: A Self-Evolving Framework for Automated Machine Learning Algorithm Discovery
Self-Augmenting Retrieval for Diffusion Language Models
Taiji: Pareto Optimal Policy Optimization with Semantics-IDs Trade-off for Industrial LLM-Enhanced Recommendation
Training-Free Lexical-Dense Fusion for Conversational-Memory Retrieval
The Loss Is Not Enough: Sampling Conditions and Inductive Bias in Contrastive Representation Learning
Argus-Retriever: Vision-LLM Late-Interaction Retrieval with Region-Aware Query-Conditioned MoE for Visual Document Retrieval
Bridging the Semantic-Collaborative Gap: An Asymmetric Graph Architecture for Cold-Start Item Recommendation
Benchmarking Open-Source Layout Detection Models for Data Snapshot Extraction from Institutional Documents
OneReason Technical Report
A Vision-language Framework for Comparative Reasoning in Radiology
MDP-GRPO: Stabilized Group Relative Policy Optimization for Multi-Constraint Instruction Following
Multi-task Learning is Not Enough: Representational Entanglement in Dual-output Second Language Speech Recognition
SkillComposer: Learning to Evolve Agent Skills for Specification and Generalization
On Advantage Estimates for Max@K Policy Gradients
Attention Calibration for Position-Fair Dense Information Retrieval
MemoryCard: Topic-Aware Multi-Modal Clue Compression for Long-Video Question Answering
Asuka-Bench: Benchmarking Code Agents on Underspecified User Intent and Multi-Round Refinement
Retrospective Harness Optimization: Improving LLM Agents via Self-Preference over Trajectory Rollouts
Better Literary Translation: A Multi-Aspect Data Generation and LLM Training Approach
МЕТОДЫ АНАЛИЗА ТЕКСТОВЫХ ДАННЫХ С ИСПОЛЬЗОВАНИЕМ МАШИННОГО ОБУЧЕНИЯ
КОНЦЕПЦИЯ ИЗУЧЕНИЯ ОБРАБОТКИ ЕСТЕСТВЕННОГО ЯЗЫКА (NLP) В ВЫСШЕЙ ШКОЛЕ
ОБРАБОТКА ТЕКСТА С ПОМОЩЬЮ НЕЙРОСЕТИ В НАПРАВЛЕНИИ АВТОРСТВА
ИСПОЛЬЗОВАНИЕ ТЕХНОЛОГИИ TEXT MINING ПРИ АВТОМАТИЧЕСКОЙ ОБРАБОТКЕ ТЕКСТА
Deep Learning for NLP
Text Summarization using Formal Argumentation
Caliper: Probing Lexical Anchors versus Causal Structure in LLMs
Dual-Stream MLP is All You Need for CTR Prediction
NLLog: Lightweight, Explainable SOC Anomaly Detection via Log-to-Language Rewriting
SearchLog: A Web Browser Extension for Capturing Search Logs in Laboratory Studies
Legal documents Text Analysis using Natural Language Processing (NLP)
Expert-Aware Causal Tracing of Factual Recall in Sparse MoE Language Models
Reasoning over Grammar: Can Synthetic Linguistic Reasoning Traces Enhance Low-Resource Machine Translation?
Backdoor Unlearning Generalization: A Path Toward the Removal of Unknown Triggers in LLMs
Exploring Adversarial Robustness and Safety Alignment in Multilingual Multi-Modal Large Language Models
Decoding Conversational AI: From Text to Context with NLP
Assisted strategic monitoring on call for tender databases using natural language processing, text mining and deep learning
Multimodal Approaches for Visually-Rich Document Type Classification: A Comparative Analysis
Harness-1: Reinforcement Learning for Search Agents with State-Externalizing Harnesses
Dynamic Spectral Denoising with Global-Context Attention for Multi-Behavior Recommendation
HERO'S JOURNEY: Testing Complex Rule Induction with Text Games
From Layers to Submodules: Rethinking Granularity in Replacement-Based LLM Compression
Trustworthy Recommendation in the Era of Large Language Models: Opportunities and Challenges
Critic-R: Improving Agentic Search using Instruction-tuned Retrievers with Natural Language Introspective Feedback
MemGraphRAG: Memory-based Multi-Agent System for Graph Retrieval-Augmented Generation
SpikeHash: Learning Binary Codes with Spiking Neural Networks for Cross-Modal Hashing Retrieval
K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts
Investigating and Alleviating Harm Amplification in LLM Interactions
ODTQA-FoRe: An Open-Domain Tabular Question Answering Dataset for Future Data Forecasting and Reasoning
On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters
SkillBrew: Multi-Objective Curation of Skill Banks for LLM Agents
Do Gender Cues Affect LLM Value Trade-offs? Evidence from a Controlled Decision Benchmark
Better with Experience: Self-Evolving LLM Agents for Evidence-Grounded Health Community Notes
When Knowledge Is Not Free: Cost-Aware Evidence Selection in Retrieval-Augmented Generation
Easier to Mislead Than to Correct: Harmful and Beneficial Revision in LLM Conformity
MobEvolve: An Agentic Self-Evolving Heuristic System for Interpretable Human Mobility Generation
When Meaning Travels: A Granular Lens on Hybrid-MoE's Role in Idiomatic Understanding for Language Models
Why Do Self-Harm Prediction Models Struggle to Generalise? Lexical and Semantic Variations in Emergency Department Triage Notes
PROTOCOL: Late Interaction Retrieval for Protein Homolog Search
Mellum2 Technical Report
Wind Turbine Maintenance Log Labelling Framework: LLM-Driven Data Correction and Enrichment via Semantic Extraction of Reliability Intelligence
Divergence Decoding: Inference-Time Unlearning via Auxiliary Models
Disagreeing Rationales: Rethinking Classification and Explainability Evaluation in Hate Speech Detection
What Gets Unmasked First? Trajectory Analysis of Diffusion Models for Graph-to-Text Generation
Language Models Learn Constructional Semantics, Not To Mention Syntax: Investigating LM Understanding of Paired-Focus Constructions
ConsisGuard: Aligning Safety Deliberation with Policy Enforcement in LLM Guardrails
A Pilot Study on Curator-Guided Multilingual Art Description for Blind and Low-Vision Audiences with Small Vision-Language Models
Beyond Static Dialogues: Benchmarking Realistic, Heterogeneous, and Evolving Long-Term Memory
KnowledgeGain: Evaluating and Optimizing Science News Generation for Reader Learning
DocRetriever: A Plug-and-Play Framework for Multimodal Document Retrieval with Comprehensive Benchmark
No More K-means:Single-Stage Sparse Coding for Efficient Multi-Vector Retrieval
LexPath: A domain-oriented multi-path framework for legal article retrieval
When Should Models Change Their Minds? Contextual Belief Management in Large Language Models
Do Language Models Track Entities Across State Changes?
ConvMemory: A Lightweight Learned Memory Reranker, a Negative Attribution Result, and a Research-Preview Conflict Editor
Generative Spatiotemporal Intent Sequence Recommendation via Implicit Reasoning in Amap
SilentRetrieval: Hijacking Retrieval-Augmented Generation via Semantically-Preserving Adversarial Data Poisoning
A Wolf in Sheep's Clothing: Targeted Routing Hijacking in Federated RAG
NLP in Customer Service
Give it Space! Explicit Disentangling of Positional and Semantic Representations in Encoders
Audio Jailbreaks in Large Audio-Language Models: Taxonomy, Attack-Defense Analysis, and Cost-Aware Evaluation
Teaching Values to Machines: Simulating Human-Like Behavior in LLMs
Token Inflation: How Dishonest Providers Can Overcharge for Large Language Model Usage
On the Practice of Scaling Search Conversion Rate Prediction
Rethinking Literature Search Evaluation: Deep Research Helps, and Human Citation Lists Are Not a Ground Truth
Surfacing Isolated Learners with Outcome-Independent Mediation of Feedback between Teachers and Students Using AI
OmniRetrieval: Unified Retrieval across Heterogeneous Knowledge Sources
SAAS: Self-Aware Reinforcement Learning for Over-Search Mitigation in Agentic Search
Metric-Dependent Annotation Saturation for Learning from Label Distributions
Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels
AgentDoG 1.5: A Lightweight and Scalable Alignment Framework for AI Agent Safety and Security
Locally Coherent, Globally Incoherent: Bounding Compositional Incoherence in Multi-Component LLM Agents
Unlocking the Working Memory of Large Language Models for Latent Reasoning
SchGen: PCB Schematic Generation with Semantic-Grounded Code Representations
LLMSurgeon: Diagnosing Data Mixture of Large Language Models
AI, Take the Wheel: What Drives Delegation and Trust in Human-Computer Cooperative Question Answering?
PrunePath: Towards Highly Structured Sparse Language Models
CIRF: Tokenizing Chain-of-Thoughts into Reusable Functional Units for Efficient Latent Reasoning in Large Language Models
Where Rollouts Begin: Low-Load, High-Leverage First-Token Diversification for RLVR
Personal Visual Memory from Explicit and Implicit Evidence
Self-Improving Language Models with Bidirectional Evolutionary Search
Descriptive Feedback on Interns’ Performance using a text mining approach
Credit-assigned Policy Gradient for Early Stage Retrieval in Two-stage Ranking
Plans for Evaluating Structured Generative Search Summaries
Uniboost: Global Coordination with Value Alignment for Fair and Efficient Traffic Allocation
EpiCurveBench: Evaluating VLMs on Epidemic Curve Digitization
Temporal Simultaneity Predicts Annotation Quality in Sentiment Corpora
Review of: "Building Urban Resilience through Mega-Events: A Systematic Review using Text Mining and Natural Language Processing (NLP)"
Search-E1: Self-Distillation Drives Self-Evolution in Search-Augmented Reasoning
Rethinking Agentic RAG: Toward LLM-Driven Logical Retrieval Beyond Embeddings
GraphReview: Scientific Paper Evaluation via LLM-Based Graph Message Passing
The Coverage Illusion: From Pre-retrieval Routing Failure to Post-retrieval Cascades in a Production RAG System
Separating Semantic Competition from Context Length in RAG Reading
Benchmarking Patent Embeddings: A Multi-Task Evaluation of 22 Models Across Retrieval, Classification, and Clustering
SemanticZip: A Pilot Framework for Lossy Text Compression with LLMs as Semantic Decompressors
Beyond Control-Flow: Integrating the Resource Perspective into Multi-Collaborative Process Modeling from Text
The Multilingual Curse at the Retrieval Layer: Evidence from Amharic
QUIET: A Multi-Blank Cascaded Story Cloze Benchmark for LLM Creative Generation Capability
PolyGnosis 2.0: Enhancing LLM Reasoning via Agentic Harness Engineering for Polymarket and OSINT Insight Extraction
Mapping the Schedule x Bit-Width Boundary in Sub-100M Quantisation-Aware Training
Triplet-Block Diffusion RWKV
SIREN: Unified Multi-Granularity Semantic Interaction for Multi-Modal Lifelong User Interest Modeling
DeGRe: Dense-supervised Generative Reranking for Recommendation
Anticipate and Learn: Unleashing Idle-Time Compute in Proactive Agents
SemBridge: Language Transfer in Sparse Encoders via Multilingual Semantic Bridges
BiRD: A Bidirectional Ranking Defense Mechanism for Retrieval Augmented Generation
SAGE: Scalable Automatic Gating Ensemble for Confident Negative Harvesting in Fraud Detection
Rethinking Contrastive Learning for Graph Collaborative Filtering: Limitations and a Simple Remedy
Layer-wise Token Compression for Efficient Document Reranking
Natural Language Processing (NLP) Methods for Determining the Style of Kazakh Text
BC Protocol: Structured Dual-Expert Dialogue for Eliciting High-Quality Chain-of-Thought Post-Training Data
RotMoLE: Enhancing Mixture of Low-Rank Experts through Rotational Gating Mechanism
PennySynth: RAG-Driven Data Synthesis for Automated Quantum Code Generation
Multilingual Phonological Feature Recognition with Self-Supervised Speech Models
Metacognition as Reward: Reinforcing LLM Reasoning via Knowledge and Regulation Signals
EquiSumm : A Gender Bias-Aware Framework for Inclusive Tweet Summarization
Articulatory strategy as a source of variation in acoustic vowel dynamics
Naturalistic measure of social norms alignment
Echoes in Filter Bubble: Diagnosing and Curing Popularity Bias in Generative Recommenders
Unlocking Biological Workflows for Robust Protein-Text Question Answering: A Dual-Dimensional RAG Framework
RAGR: Review-Augmented Generative Recommendation
Dual-Diffusional Generative Fashion Recommendation
SomaliWeb v1: A Quality-Filtered Somali Web Corpus with a Matched Tokenizer and a Public Language-Identification Benchmark
RCTEA: Richness-guided Co-training for Temporal Entity Alignment
Efficient Table QA via TableGrid Navigation and Progressive Inference Prompting
From Volume to Value: Preference-Aligned Memory Construction for On-Device RAG
As X, Do Y: How Persona and Task Combine in Instruction-Tuned LLMs
When Symptoms Are Not Enough: Evidence-Weighting Patterns in Large Language Model Psychiatric Screening
Same Model, Different Weakness: How Language and Modality Reshape the Jailbreak Attack Surface in Frontier MLLMs
What Does the Server See? Understanding Privacy Leakage from Large Language Models in Split Inference
The Evolution of Natural Language Processing
A Reproducible Universal Dependencies-Style Pipeline for Katharevousa Greek Parliamentary Text
Memorization Dynamics of Fill-in-the-Middle Pretraining
A Proactive Multi-Agent Dialogue Framework for Assessing Social Language Disorder Traits in Autism
The Deterministic Horizon: Impossibility Results as Design Specifications for Trustworthy AI Systems
Strong Teacher Not Needed? On Distillation in LLM Pretraining
Multilingual Knowledge Transfer under Data Constraints via Lexical Interventions
ETCHR: Editing To Clarify and Harness Reasoning
SkillOpt: Executive Strategy for Self-Evolving Agent Skills
UNDERSTANDING NATURAL LANGUAGE PROCESSING (NLP) TECHNIQUES: FROM TEXT ANALYSIS TO LANGUAGE GENERATION
Natural Language Processing for Social Media Data Mining
Natural Language Processing (NLP) Methods for Determining the Style of Kazakh Text
Temporal Decay of Co-Citation Predictability: A 20-Year Statute Retrieval Benchmark from 396M Ukrainian Court Citations
News Signals: An NLP Library for Text and Time Series
Advancing NLP models with strategic text augmentation: A comprehensive study of augmentation methods and curriculum strategies
Reducing Political Manipulation with Consistency Training
Evaluating Commercial AI Chatbots as News Intermediaries
Tokenisation via Convex Relaxations
Linguistic Computing with UNIX Tools
Evolving Explanatory Novel Patterns for Semantically-Based Text Mining
Search-E1: Self-Distillation Drives Self-Evolution in Search-Augmented Reasoning
One prompt is not enough: Instruction Sensitivity Undermines Embedding Model Evaluation
Diversed Model Discovery via Structured Table Discovery
Understanding Wacky Weights: A Dissection of SPLADE's Learned Term Importance
Divergence Meets Consensus: A Multi-Source Negative Sampling Framework for Sequential Recommendation
Auditing Privacy in Multi-Tenant RAG under Account Collusion
X-SYNTH: Beyond Retrieval -- Enterprise Context Synthesis from Observed Digital Human Attention
SynAE: A Framework for Measuring the Quality of Synthetic Data for Tool-Calling Agent Evaluations
LANG: Reinforcement Learning for Multilingual Reasoning with Language-Adaptive Hint Guidance
Beyond Temperature: Hyperfitting as a Late-Stage Geometric Expansion
A Tutorial on Diffusion Theory: From Differential Equations to Diffusion Models
Pattern-and-root inflectional morphology: the Arabic broken plural
TransitLM: A Large-Scale Dataset and Benchmark for Map-Free Transit Route Generation
Modeling Pathology-Like Behavioral Patterns in Language Models Through Behavioral Fine-Tuning
Boundary-targeted Membership Inference Attacks on Safety Classifiers
Analysis of Research Trends on startup Business Performance Using Natural Language Processing(NLP)-Based Text Mining : Focusing on RISS DB
FlyRoute: Self-Evolving Agent Profiling via Data Flywheel for Adaptive Task Routing
Hy-MT2: A Family of Fast, Efficient and Powerful Multilingual Translation Models in the Wild
Faithful-MR1: Faithful Multimodal Reasoning via Anchoring and Reinforcing Visual Attention
From Reasoning Chains to Verifiable Subproblems: Curriculum Reinforcement Learning Enables Credit Assignment for LLM Reasoning
Building Urban Resilience through Mega-Events: A Systematic Review using Text Mining and Natural Language Processing (NLP)
Text Mining
Review of: "Building Urban Resilience through Mega-Events: A Systematic Review using Text Mining and Natural Language Processing (NLP)"
Biomedical Natural Language Processing and Text Mining
Text Mining
Review of: "Building Urban Resilience through Mega-Events: A Systematic Review using Text Mining and Natural Language Processing (NLP)"
MARQUIS: A Three-Stage Pipeline for Video Retrieval-Augmented Generation
Uncertainty-Calibrated Recommendations for Low-Active Users
Accelerating AI-Powered Research: The PuppyChatter Framework for Usable and Flexible Tooling
DADF: A Distribution-Aware Debiasing Framework for Watch-Time Regression in Recommender Systems
KoRe: Compact Knowledge Representations for Large Language Models
ClinSeekAgent: Automating Multimodal Evidence Seeking for Agentic Clinical Reasoning
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models
TIDE: Efficient and Lossless MoE Diffusion LLM Inference with I/O-aware Expert Offload
The 99% Success Paradox: When Near-Perfect Retrieval Equals Random Selection
Towards Self-Evolving Agentic Literature Retrieval
Efficient Generative Retrieval for E-commerce Search with Semantic Cluster IDs and Expert-Guided RL
Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture
Fairness-Aware Retrieval Optimization for Retrieval-Augmented Generation
Generative Long-term User Interest Modeling for Click-Through Rate Prediction
LERA: LLM-Enhanced RAG for Ad Auction in Generative Chatbots
CLIF: Concept-Level Influence Functions for Transparent Bottleneck Models
Are Tools Always Beneficial? Learning to Invoke Tools Adaptively for Dual-Mode Multimodal LLM Reasoning
Where Does Authorship Signal Emerge in Encoder-Based Language Models?
SENTIMENT ANALYSIS OF STUDENTS’ FEEDBACK ON INSTITUTIONAL FACILITIES USING TEXT-BASED CLASSIFICATION AND NATURAL LANGUAGE PROCESSING (NLP)
🚀 NLP Colloquium
How Do Document Parsers Break? Auditing Structural Vulnerability in Document Intelligence
A Multi-Agent Framework for Feature-Constrained Difficulty Control in Reading Comprehension Item Generation
STAR-PólyaMath: Multi-Agent Reasoning under Persistent Meta-Strategic Supervision
HalluWorld: A Controlled Benchmark for Hallucination via Reference World Models
ANALYSIS OF EXTENDED REALITY PUBLCATONS IN INFORMATION SYSTEMS RESEARCH AREA THROUGH TEXT MINING AND NATURAL LANGUAGE PROCESSING (NLP) TECHNIQUES
Natural Language Processing (NLP) Trends
Harnessing Natural Language Processing (NLP) and Generative AI Techniques for Social Media Sentiment Analysis with Text Classification
Deep Learning for NLP
Biomedical Text Mining: Applicability of Machine Learning-based Natural Language Processing in Medical Database
Text Mining and Natural Language Processing
Agentic Chunking and Bayesian De-chunking of AI Generated Fuzzy Cognitive Maps: A Model of the Thucydides Trap
A Pilot Benchmark for NL-to-FOL Translation in Planetary Exploration
Prompt Compression in Diffusion Large Language Models: Evaluating LLMLingua-2 on LLaDA
Universal Adversarial Triggers
Task-Aware Automated User Profile Generation for Recommendation Simulation Using Large Language Models
Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation
Ancient Greek to Modern Greek Machine Translation: A Novel Benchmark and Fine-Tuning Experiments on LLMs and NMT Models
Easier to Judge than to Find: Predicting In-Context Learning Success for Demonstration Selection
Continuous Diffusion Scales Competitively with Discrete Diffusion for Language
STT-Arena: A More Realistic Environment for Tool-Using with Spatio-Temporal Dynamics
Analysis of Aphasia Research Trends Using Natural Language Processing(NLP) Based Text Mining
Natural Language Processing: An Introduction
Review of: "Building Urban Resilience through Mega-Events: A Systematic Review using Text Mining and Natural Language Processing (NLP)"
Review of: "Building Urban Resilience through Mega-Events: A Systematic Review using Text Mining and Natural Language Processing (NLP)"
Review of: "Building Urban Resilience through Mega-Events: A Systematic Review using Text Mining and Natural Language Processing (NLP)"
Review of: "Building Urban Resilience through Mega-Events: A Systematic Review using Text Mining and Natural Language Processing (NLP)"
Ontology for Policing: Conceptual Knowledge Learning for Semantic Understanding and Reasoning in Law Enforcement Reports
Defining Cultural Capabilities for AI Evaluation: A Taxonomy Grounded in Intercultural Communication Theory
Can Vision Language Models Be Adaptive in Mathematics Education? A Learner Model-based Rubric Study
Judge Circuits
Text Mining for Biomedicine
iSAI-NLP Committee
iSAI-NLP Committee
Chapter 5. Text mining
NLP Applications – Active Usage
Ascend-RaBitQ: Heterogeneous NPU-CPU Acceleration of Billion-Scale Similarity Search with 1-bit Quantization
MERVIN: A Unified Framework for Multimodal Event Retrieval in Vietnamese News Videos
paper.json: A Coordination Convention for LLM-Agent-Actionable Papers
Argus: Evidence Assembly for Scalable Deep Research Agents
Toward LLMs Beyond English-Centric Development
Evaluating Chinese Ambiguity Understanding in Large Language Models
Dynamic Chunking for Diffusion Language Models
VCG-Bench: Towards A Unified Visual-Centric Benchmark for Structured Generation and Editing
A Comprehensive Review of Natural Language Processing Methods for Text Mining Applications
Extracting Product Features and Opinions from Reviews
iSAI-NLP Committee
NLP (Natural Language Processing) and NLP (Natural Language Programming) Using Decision Making Test and Evaluation Laboratory (DEMATEL) Method
The Scientific Contribution Graph: Automated Literature-based Technological Roadmapping at Scale
Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance
Small, Private Language Models as Teammates for Educational Assessment Design
COTCAgent: Preventive Consultation via Probabilistic Chain-of-Thought Completion
Natural Language Processing (NLP): A Comparative Study on Applications of GNNs for the Tasks Like Text Classification, Semantic Role Labelling, and Abstract Meaning Representation Parsing
Universal Segmentation of Text with the Sumo Formalism
NLP in Online Reviews
Automatic Evaluation of Ontologies
Building Large Resources for Text Mining: The Leipzig Corpora Collection
NLP in Virtual Assistants
Nexus : An Agentic Framework for Time Series Forecasting
Agentic Recommender System with Hierarchical Belief-State Memory
Knowledge Beyond Language: Bridging the Gap in Multilingual Machine Unlearning Evaluation
SWE-Chain: Benchmarking Coding Agents on Chained Release-Level Package Upgrades
NLP applications
8th International Conference on Natural Language Processing (NLP 2019)
NLP Applications - Developing Usage
Text Mining and Emotion Classification on Monkeypox Twitter Dataset: A Deep Learning-Natural Language Processing (NLP) Approach
AN APPROACH TO SEMANTIC EDUCATIONAL CONTENT MINING USING NATURAL LANGUAGE PROCESSING (NLP)
BioNLP: Biomedical Text Mining K. Bretonnel Cohen
Mining Profiles and Definitions with Natural Language Processing
Natural Language Processing Supporting Interoperability in Healthcare