Overview
The “Agentic and Generative AI for E-Commerce” workshop explores the rapidly evolving intersection of recommender systems, generative AI, and agentic AI in online retail. As AI systems evolve from passive generators to autonomous agents capable of planning, reasoning, and taking actions, e-commerce stands at the forefront of this transformation. The objective of this workshop is to foster discussions on how agentic AI systems — autonomous agents that can browse, compare, negotiate, and purchase on behalf of users — alongside generative models ranging from large language models (LLMs) to diffusion-based techniques, can transform personalization, product recommendations, content creation, and user engagement in e-commerce platforms. Through this workshop, we seek to highlight novel research, industry applications, and emerging trends that can enhance the capabilities of modern recommender systems.
E-commerce companies face challenges such as lack of quality content, subpar user experience, and sparse datasets. Generative and agentic AI offer significant potential to address these — from generating product content to deploying autonomous shopping assistants for end-to-end purchase workflows. Yet, scaling these technologies presents challenges including hallucination, excessive costs, latency, and ensuring safe autonomous agent behavior.
Call for Papers
We welcome papers that leverage Agentic and Generative Artificial Intelligence (Gen AI) in e-commerce. Detailed topics are mentioned in CFP. Papers can be submitted at Easychair. Accepted papers will be published as open-access workshop proceedings on this workshop website (as we did for the previous edition), rather than via CEUR-WS this year. See the CFP page for details.
Important Dates
- Call for Papers publication: April 21, 2026
- Paper submission deadline: July 20, 2026
- Reviewer deadline: August 7, 2026
- Author notification: August 14, 2026
- Camera-ready version deadline: September 1, 2026
- Workshop: September 28, 2026 (13:30–17:30)
Schedule
We have a half-day program at Minneapolis, Minnesota, USA.
| Time | Agenda |
|---|---|
| 1:30–1:40 PM | Registration and Welcome |
| 1:40–2:20 PM | Keynote by Patrick Jordan (Microsoft): When Everyone Has an Agent: From Better Decisions to Better Markets |
| 2:20–3:00 PM | Keynote by Akshay Soni (Shopify): Foundation Models for Agentic and Counterfactual Decision Support in E-Commerce |
| 3:00–3:30 PM | Coffee Break |
| 3:30–4:10 PM | Keynote by Shengbo Guo (Meta): LLM Ranking in Facebook Verticals: from Content-Based LLM Ranking to a Unified Generative & Ranking Recommender |
| 4:10–4:25 PM | Paper Presentation 1: The Disconnect Between Better Descriptive Reasoning Trace Quality and Recommendation Effectiveness — Gustavo Penha |
| 4:25–4:40 PM | Paper Presentation 2: Ad Insertion in LLM-Generated Responses — Shengwei Xu |
| 4:40–5:30 PM | Poster Session |
Keynote Speakers
Patrick Jordan
Microsoft
When Everyone Has an Agent: From Better Decisions to Better Markets
Akshay Soni
Shopify
Foundation Models for Agentic and Counterfactual Decision Support in E-Commerce
Shengbo Guo
Meta
LLM Ranking in Facebook Verticals: from Content-Based LLM Ranking to a Unified Generative & Ranking Recommender
Accepted Papers
- EvoRank: LLM-Guided Evolution of Multi-Objective Learning-to-Rank Pipelines
Rayhan Patel, Shabaz PatelAbstractAbstract: We present EvoRank, an open autonomous ranking engineer: an LLM-guided evolutionary loop that discovers complete Learning-to-Rank pipelines (features, models, losses, ensembles) for multi-objective e-commerce search. On the Expedia ICDM 2013 dataset, with relevance, conversion, and revenue as competing objectives, three independent runs each converge within 50 iterations (about ten dollars) on interpretable pipelines that beat an Optuna-tuned LambdaMART on 60k held-out queries, an advantage that persists at full data scale and places in the top 6 percent of the original competition. A first campaign, evolving only training objectives, builds the central design rule: it appeared to work on its selection fold (the small dataset it uses to pick winners) while a transfer audit, re-scoring winners on held-out data, showed the gains were almost entirely fitness noise (the randomness of its own scoring), and neither seeded domain knowledge nor richer diagnostic feedback changed what transferred. The deciding quantity is measurable in advance: search-space headroom relative to fitness noise. We package this as a headroom gate that predicts, before any LLM spend, whether the loop will pay off, and we release the system, the auditing tools, and a catalog of failure modes with their guardrails, so teams can apply the procedure to their own ranking stacks.PDF Code - SR-Agent: An Experience-Driven Agentic Framework for Post-Ranking Strategy Refinement in E-Commerce Recommendation
Hanchen Yang, Kaiwen Yang, Junpeng Zhuang, Yang He, Keting Cen, Bochao Liu, Zhongbo Sun, An Liu, Zhongteng Han, Chenyi LeiAbstractAbstract: User experience is a first-class objective in industrial e-commerce recommender systems (RS). Post-ranking strategies, which govern diversity, similarity, and exposure over a ranked list, are widely deployed in industrial RS for their simplicity and low serving cost. However, as the online recommendation environment evolves continuously, these statically configured strategies gradually become stale, thereby degrading the user experience. Refining them typically relies on manual inspection, diagnosis, and updates, making it slow, costly, and difficult to scale or reuse. Although recent LLM-based agents (e.g., RecUserSim, SimUSER, and Self-EvolveRec) offer promising directions, none of them close the full loop of automated, self-evolving strategy refinement. To bridge this gap, we introduce SR-Agent, which, to the best of our knowledge, is the first agentic framework deployed to refine post-ranking strategies in industrial RS. SR-Agent unifies three components: (i) a UserSim agent that applies inspection skills to surface user-perceived bad cases; (ii) an Analysis agent that consolidates recurring bad cases into structured, reusable diagnoses; and (iii) a constrained Strategy Refinement Harness that maps diagnoses to typed and bounded actions, gated by a four-stage reward pipeline with reversible rollback. Deployed on the Kuaishou e-commerce platform, SR-Agent continuously runs this refinement loop and, in a one-month online A/B test, increases order volume by 0.71%, browsing depth by 0.34%, and clicked-category diversity by 0.48%, while markedly shortening the refinement cycle and lowering operational cost.PDF Code - MASLOW: Multi-Agent Synthetic Data Generation for Underspecified, Low-Resource E-Commerce Classification
Ling Jiang, Shaobai Jiang, Ankur Gupta, Adalie Palma, Ali Marjaninejad, Xiaoyu Chu, Prithvi SenAbstractAbstract: Training classifiers for e-commerce catalog management is bottlenecked by two simultaneous scarcities: underspecified task specifications (a bare URL or one-sentence description) and few or no labeled examples. We present MASLOW, to our knowledge the first multi-agent synthetic data pipeline that handles the full journey from underspecified inputs to labeled training data without requiring clean class labels, curated seeds, or task-specific corpora. Four specialized agents progressively enrich specifications, identify decision boundaries, generate diverse examples including hard negatives, and filter for quality. We evaluate MASLOW on 1,925 low-resource product-matching tasks and 40 unseen product classification tasks, two orders of magnitude beyond typical synthetic data studies. Ablation reveals that upstream context enrichment has outsized impact: removing the context-enrichment stage causes twice the performance drop (−5.8% F1) of removing any other agent, suggesting practitioners should prioritize input quality over generation architecture. Synthetic data complements human labels (+11.1% F1 on the scarcest tasks) and enables cold-start training from specifications alone. For In-Context Learning, task-aligned feedback (MASLOW+AFB) that regenerates examples receiving uncertain downstream predictions recovers quality that naive feedback loses. At 47× lower cost and 31× lower latency than human annotation, MASLOW processes 1,925 tasks in approximately one hour.PDF Code - AURA: Agentic Diagnosis and Refinement for Production Recommender Systems at Scale
SungGeun Kim, Abhinav Narain, Daniel NemirovskyAbstractAbstract: How and why does a recommender system fail the users it serves? Oftentimes in recommender systems, practitioners are left to improve their algorithms based on a combination of feedback from stakeholder teams, domain expertise, and insights from data analyses. Yet, the nuances of how and where the recommendations are performing well or poorly for the end users are difficult to discern from aggregate quantitative metrics. Whereas these metrics can provide a high-level and incomplete picture, further granularity into the quality of recommendations and their patterns requires reasoning with domain understanding and objectivity, at scale. In this paper, we contemplate this complex conundrum and describe a method and implementation that utilizes the latest AI agentic advances to provide actionable diagnoses and improvements for the production recommender systems. We present AURA (Agentic Understanding and Refinement of recommender Algorithms), an end-to-end agentic system that performs qualitative evaluation at scale and can then generate improvements to our algorithms at the code level. Specialized agents read production engagement logs, from thousands of sessions to millions, and surface patterns and examples of how the recommender fails real users. The next step uses those diagnoses as well as context about the recommender's own code, data, and training pipeline to propose and implement refinements grounded in that codebase. We report the system design, initial tests on production data from two large consumer platforms at a major media-streaming company, safeguards, operational learnings, and early results toward a self-improving recommender system. Finally, the diagnostic gap AURA closes is not specific to streaming. The architecture is built to transfer and every domain-specific element enters through the configuration layer that already ported it between our two platforms. We map it concretely to e-commerce and online-retail recommendation.PDF Code - The Disconnect Between Better Descriptive Reasoning Trace Quality and Recommendation Effectiveness
Gustavo Penha, Juan Elenter, Claudia Hauff, Hugues Bouchard, Paul Bennett, Mounia LalmasAbstractAbstract: Recent work has focused on improving explicit natural-language descriptive reasoning traces for generative recommendation. This includes systems that augment semantic ID (SID) prediction with chain-of-thought reasoning. However, because SIDs are opaque learned identifiers rather than natural language, they require costly alignment before an LLM can reason over them. This provides a controlled experimental setting in which both item representation (Title vs. SID) and semantic grounding (minimal vs. extensive SID alignment) can be varied independently. We therefore present the first controlled comparison of descriptive reasoning trace quality across semantic IDs and natural-language titles in a 2 × 2 factorial study on three Amazon product domains using a shared Qwen3-1.7B backbone. We find that introducing explicit descriptive reasoning traces reduces traditional offline recommendation effectiveness under standard SFT and RL training, even though natural language titles produce substantially more grounded and interpretable traces. Extensive SID alignment improves descriptive trace quality but not traditional offline recommendation effectiveness, while a richer reward signal partially recovers performance. Overall, our results show that improving descriptive reasoning trace quality is not, by itself, sufficient to consistently improve traditional offline recommendation effectiveness under the training objectives and evaluation protocols studied here.PDF Code - Efficient Clustering with Quality Guardrails for LLM-based Recommender Systems at Industry Scale
Longshaokan Wang, Wai Tsang Keung, Punit Ghodasara, Roman Wang, Ali Dashti, Francesc Moreno-NoguerAbstractAbstract: LLMs can be prohibitively expensive and slow to run at scale, especially for applications that invoke an LLM per sample over millions of inputs. A natural way to scale such applications is to cluster the inputs, run the LLM only on cluster representatives, and propagate the LLM outputs to other cluster members. However, in this setup, the LLM outputs a cluster member receives are only as good as the member's match to the cluster representative. Off-the-shelf clustering methods optimize an aggregate objective over all samples, targeting only average-case quality without offering per-sample quality guardrails. As a result, cluster members can be assigned to poorly-matched representatives, and the inherited LLM outputs — though appropriate for the representative — may be irrelevant or even unsafe for the member. For example, a parent of a toddler grouped with parents of older children could receive age-inappropriate recommendations. Furthermore, most clustering methods scale poorly to millions of inputs in runtime and memory, limiting their applicability at industry scale. We propose a scalable two-stage clustering algorithm with provable per-sample quality guardrails: every sample is guaranteed to share a user-specified minimal similarity and exact attribute match with its cluster representative. The algorithm first generates initial clusters with Mini-batch K-Means, then greedily selects representatives within each initial cluster to satisfy the guardrails. We provide theoretical guarantees, complexity analysis, and benchmarks against common clustering methods on internal and public datasets. We show that our method not only delivers per-sample guardrails but also runs substantially faster and scales to data sizes where most standard methods become intractable. We demonstrate our method's impact in a real-world deployment clustering 38 million customers. We reduce downstream LLM cost and runtime by 50× while preserving personalization, unblocking the launch of a persona-based recommender system that delivers significant gains in revenue and engagement in an A/B test.PDF Code - Ad Insertion in LLM-Generated Responses
Shengwei Xu, Zhaohua Chen, Xiaotie Deng, Zhiyi Huang, Grant SchoenebeckAbstractAbstract: Sustainable monetization of large language models (LLMs) remains a critical open challenge. Traditional search advertising, which relies on static keywords, fails to capture the fleeting, context-dependent user intent—the specific information, goods, or services a user seeks—embedded in conversational flows. Beyond the standard goal of social welfare maximization, effective LLM advertising requires contextual coherence (aligning ads semantically with transient user intent), computational efficiency (avoiding user-facing latency), and adherence to ethical and regulatory standards, including privacy preservation and explicit ad disclosure. Although recent solutions have explored bidding at the token and query levels, neither category holistically satisfies these constraints. We propose a framework that resolves these tensions through two decoupling strategies. First, we decouple ad insertion from response generation to facilitate pre-screening and explicit disclosure. Second, we decouple bidding from specific user queries by using "genres" (high-level semantic clusters) as a proxy. This allows advertisers to bid on stable categories rather than sensitive real-time responses, reducing computational burden and privacy risks. Applying the VCG auction mechanism to this genre-based framework provides approximate guarantees for dominant-strategy incentive compatibility (DSIC), individual rationality (IR), and social welfare. In synthetic experiments with 105 advertisers and 100 candidate slots, VCG clears in approximately 1.25 seconds on a consumer-grade laptop. Finally, we introduce an "LLM-as-a-Judge" metric for estimating contextual coherence. Its predictions correlate with mean human ratings at Spearman's 𝜌 ≈ 0.66 and have a higher correlation with the leave-one-out group mean than 29 of 36 individual raters (80.6%).PDF Code - Benchmarking Models for Conversational E-Commerce: A Reproducible Evaluation Framework
Yuri Vorontsov, Diogo S. Carvalho, Anastasia Vorontsov, Anna Platonova, Vladimir Gorovoy, Ilya Briskin, Felix Tseitlin, Senka Krivic, Salman AhmadAbstractAbstract: We present a reproducible framework for benchmarking large language models as multi-turn conversational e-commerce agents across the full shopping loop of search, clarification, recommendation, and add to cart. We demonstrate the framework in a first-in-class 12-model study, complemented by a smaller cross-category evaluation. The framework comprises product catalog ingestion, scenario generation grounded in real consumer search queries, a model-breaking scenario selection filter, LLM-driven customer simulation, and a five-rubric LLM judge covering intent understanding, clarification behaviour, recommendation quality, add-to-cart correctness, and grounding fidelity. We apply the framework to 12 models on a primary jeans case study (1,200 scored conversations across five runs and 20 scenarios), using it as a worked example of the shopping loop; these scores are not intended to generalise across e-commerce. We also conduct a cross-category study on three additional product domains under a different scenario protocol. Rank order on that smaller study is directional only (𝜌 = 0.40–0.70, 𝑛 = 5). On the primary study, grounding fidelity is the main separator between score bands, Clarification is the least saturated rubric among ranks 1–7, and the highest-scoring models form overlapping bootstrap confidence bands once run-to-run variance is accounted for. These findings rest on a judge-reliability programme: a seven-judge cross-family panel with super-judge arbitration that reduces judge swing from 24.5% to 9.2%, validated against a preliminary human gold calibration (103 rubric cells from a single rater; 11 pass / 92 fail). GPT-4.1 alone is too lenient (𝜅 = −0.13); the reliable stack reaches 88.3% accuracy against human labels, below an always-fail dummy at 89.3%, so Cohen's 𝜅 (0.273) is the more informative agreement figure. Benchmark code and data are at https://github.com/rezolved/conversational-commerce-benchmark-framework.PDF Code - Domain-Adaptive Contrastive Embeddings for Frugal Skill Routing in Agentic E-Commerce Evaluation
Yue Yu, Meisam Hejazinia, Mark Martinov Kirichev, Siamak RajabiAbstractAbstract: Large-language-model (LLM) agents increasingly operate and evaluate e-commerce workflows, and a recurring building block is a skill router: a semantic-retrieval layer that maps a free-form question to the correct skill in a catalog. The same layer lets an LLM-as-judge evaluator decide which skill should have answered and whether the catalog covers the question at all. In practice it runs on general-purpose embedders that mishandle the dense, abbreviation-heavy vocabulary of retail supply-chain operations, and the reflexive fixes, a larger general embedder or a frontier LLM prompted to route, are metered per call. We instead adapt a small open-source encoder to the domain and deliver the adaptation in the two forms a deployed system needs. The static recipe fully fine-tunes a 22M-parameter MiniLM-class encoder on typed contrastive pairs under a leakage-free skill-holdout split, lifting overall routing Hit@1 from 0.474 (a production 1024-d general embedder) to 0.656, and in-distribution Hit@1 from 0.393 to 0.746, over a compact 384-d index at zero marginal inference cost. The continual-learning recipe keeps that encoder current as the catalog grows from 23 to 43 skills: warm-starting plus a 30% replay sample holds overall Hit@1 near 0.76, matching a full retrain at roughly half the compute while avoiding catastrophic forgetting. We state the main caveat plainly: the gains concentrate on skills seen in training, and on genuinely unseen skills the general embedder still leads. A score-fusion ablation tops every arm but reintroduces a per-query general-embedder call, so we keep the standalone specialist central and point to a lightweight learned gate as future work. For the retrieval substrate of agentic e-commerce, a small specialized embedder is both more accurate in domain and far cheaper than a larger general model or a frontier LLM.PDF Code - What Matters for LLM Persona Simulation: Signal Form in Cold-Start E-Commerce Advertising
Jiwoo Choi, In Ik Lee, Woo Sung Seo, Minho Lee, Yun Young ChoiAbstractAbstract: Estimating advertising performance in cold-start settings is challenging because historical logs are unavailable before launch. Large language models (LLMs) offer a potential alternative by simulating persona-conditioned user responses, but it remains unclear which design choices actually drive LLM-based persona simulation. Using real Google Ads e-commerce campaign data across five demographic segments, we identify three signal-form dimensions that affect simulation quality: (i) a demographic anchor without noisy rich attributes, (ii) a relative qualitative banner-blindness signal rather than quantitative anchors, and (iii) prompt phrasing. On this campaign, our method attains an MAE of 24.9 and a MAPE of 26.1%, a 51.3% lower MAE than the best tested non-LLM baseline and 77.3% lower than an LLM configuration using explicit CTR anchors. Systematic ablations show that minimal demographic descriptions achieve lower mean error than richer persona configurations, while relative qualitative cues substantially outperform the tested quantitative anchors. We additionally compare these configurations with direct-prior and rule-based baselines and with alternative LLM configurations.PDF Code - Beyond Ranking Accuracy: Evaluating LLM-Cited Feature Rationales for Next Basket Repurchase Recommendation
Yanan Cao, Anay Dombe, Murali Mohana Krishna Dandu, Shreeranjani Srirangamsridharan, Sinduja Subramaniam, Yogananth Mahalingam, Evren Korpeoglu, Kannan AchanAbstractAbstract: Next-basket repurchase recommendation is commonly formulated as a ranking task: given a customer's purchase history, the system ranks previously purchased items that may be needed again. In production settings, however, ranking accuracy is only one component of recommendation quality. Customers may also benefit from concise evidence about why an item is recommended now. Large language models (LLMs) offer a potential mechanism for surfacing such evidence by producing feature-based, human-readable rationales grounded in interpretable behavioral signals rather than ranking scores alone. We construct interpretable repurchase features spanning cadence, frequency, recency, user behavior, and item popularity, and evaluate LLMs on two public grocery datasets and one proprietary retail dataset. We investigate (1) whether off-the-shelf LLMs can use these features as next-basket scorers relative to personal-frequency and supervised rankers, and (2) whether the features cited by LLMs as rationales carry outcome-grounded ranking signal. For the latter, we compare LLM-cited features with model-specific attribution methods under a cross-model feature-masking protocol that measures ranking-quality degradation after masking selected features. Our results provide a nuanced view of LLM scorers for repurchase recommendation. LLM scores are not competitive with supervised rankers, suggesting that off-the-shelf LLMs should not be used directly as standalone repurchase recommenders. However, changes in prompt and evidence representation can yield stronger outcome-grounded feature-masking results in some settings even when ranking performance does not improve; this pattern is dataset-dependent and does not consistently match model-specific attribution baselines. These findings suggest a practical role for LLMs as validated explanation components: not as primary next-basket rankers, but as tools for surfacing behavioral evidence whose outcome-grounded relevance should be evaluated separately from ranking accuracy.PDF Code - Distill Globally, Adapt Locally: Reasoning Distillation and Product-Type Test-Time Training for Scalable Trade-Up Recommendation
Siliang Liu, Mohammad Ghasemi, Sapan Patel, Amin Banitalebi-DehkordiAbstractAbstract: Trade-up recommendation aims to identify higher-quality alternatives that preserve a customer's purchase intent while offering upgraded benefits through improved formulation, certifications, or brand positioning. Although large language models (LLMs) can reason about these subtle distinctions, applying them directly to hundreds of millions of product pairs in e-commerce stores is operationally impractical. We introduce a two-level framework that distills LLM-derived reasoning into an efficient non-generative student and subsequently adapts its decision boundary to product-type-specific trade-up criteria. At Level 1, a retrieval-augmented few-shot LLM teacher generates both structured relation labels and natural-language rationales. These rationales are encoded and transferred to a compact, non-generative embedding-pair classifier through alignment and contrastive objectives. At inference, the student consumes only two precomputed 768-dimensional product embeddings, requiring neither LLM calls nor text generation. On a fixed human-annotated benchmark of 8,352 pairs, a 15.5M-parameter four-class reasoning-distilled student achieves an AUC of 0.924 (95% CI [0.918, 0.929]), improving over the four-class label-only student with the same shallow architecture (AUC 0.912); rationale supervision provides little benefit when the same task is collapsed to binary labels. At Level 2, we introduce product-type test-time training (PT-TTT), which uses few-shot demonstrations as gradient-based supervision to optimize lightweight category-specific adapters over the frozen student model. PT-TTT improves AUC from 0.924 to 0.941 and average precision from 0.920 to 0.940 without serving-time LLM inference. On a 100K-pair proxy catalog, inference with the distilled student on a single eight-GPU machine is approximately 5,000× faster and has an estimated cost approximately 10,000× lower than direct LLM inference on the same workload.PDF Code - DENIM: Ablating Catalog Access and Visual Grounding in an Agentic Fashion Shopping Assistant
Alessandro Francesco Maria Martina, Cataldo Musto, Marco de Gemmis, Pasquale Lops, Giovanni SemeraroAbstractAbstract: Large Language Models now power e-commerce shopping agents that converse with customers, call retrieval tools, and recommend from live product catalogs. Building one means deciding how the product catalog reaches the agent: what it can query or read at inference time, and in what representation. We isolate two of these decisions: what kind of catalog access the agent needs, and in what form visual item information should reach it. To answer both questions, we first implemented an agentic conversational recommender for fashion e-commerce. It is designed as a single orchestrated system with tool-augmented retrieval over a large fashion catalog. We then ran controlled single-change ablations on it, evaluated with a leakage-controlled simulated shopper in a cold-start setting. First, structured catalog access is the dominant factor: removing it costs 31 points of success rate, and most of that value (27 points) comes from a catalog inspection tool that lets the agent see what the catalog contains and which attribute values are valid before filtering, a capability that recent systems include in various forms but whose effect on retrieval had not been isolated. Second, what matters about visual information is the form in which it reaches the agent: injecting item images at runtime adds no measurable success over offline-extracted visual attributes, while removing those attributes' structured rendering costs a further 11 points even though the same information remains available in prose descriptions. At matched item representation, the orchestrated pipeline also outperforms ReAct and Reflexion-style baselines built on the same backbone and tools by 12–20 points, sustaining preference elicitation where they stall. All findings replicate in direction across two simulator LLM families, one of which shares the recommender's own model family. We distill the results into practical guidance for building e-commerce shopping agents.PDF Code
Organizers
Mansi Mane
Walmart
Neeti Narayan
Amazon
Djordje Gligorijevic
Meta
Dingxian Wang
Upwork
Topojoy Biswas
Walmart
Claudio Pomo
Politecnico di Bari
Contact
For any questions, please email genai-ecommerce@googlegroups.com.