Overview

The “Agentic and Generative AI for E-Commerce” workshop explores the rapidly evolving intersection of recommender systems, generative AI, and agentic AI in online retail. As AI systems evolve from passive generators to autonomous agents capable of planning, reasoning, and taking actions, e-commerce stands at the forefront of this transformation. The objective of this workshop is to foster discussions on how agentic AI systems — autonomous agents that can browse, compare, negotiate, and purchase on behalf of users — alongside generative models ranging from large language models (LLMs) to diffusion-based techniques, can transform personalization, product recommendations, content creation, and user engagement in e-commerce platforms. Through this workshop, we seek to highlight novel research, industry applications, and emerging trends that can enhance the capabilities of modern recommender systems.

E-commerce companies face challenges such as lack of quality content, subpar user experience, and sparse datasets. Generative and agentic AI offer significant potential to address these — from generating product content to deploying autonomous shopping assistants for end-to-end purchase workflows. Yet, scaling these technologies presents challenges including hallucination, excessive costs, latency, and ensuring safe autonomous agent behavior.

Call for Papers

We welcome papers that leverage Agentic and Generative Artificial Intelligence (Gen AI) in e-commerce. Detailed topics are mentioned in CFP. Papers can be submitted at Easychair. Accepted papers will be published as open-access workshop proceedings on this workshop website (as we did for the previous edition), rather than via CEUR-WS this year. See the CFP page for details.

Important Dates

  • Call for Papers publication: April 21, 2026
  • Paper submission deadline: July 20, 2026
  • Reviewer deadline: August 7, 2026
  • Author notification: August 14, 2026
  • Camera-ready version deadline: September 1, 2026
  • Workshop: September 28, 2026 (13:30–17:30)

Schedule

We have a half-day program at Minneapolis, Minnesota, USA.

Time Agenda
1:30–1:40 PM Registration and Welcome
1:40–2:20 PM Keynote by Patrick Jordan (Microsoft): When Everyone Has an Agent: From Better Decisions to Better Markets
2:20–3:00 PM Keynote by Akshay Soni (Shopify): Foundation Models for Agentic and Counterfactual Decision Support in E-Commerce
3:00–3:30 PM Coffee Break
3:30–4:10 PM Keynote by Shengbo Guo (Meta): LLM Ranking in Facebook Verticals: from Content-Based LLM Ranking to a Unified Generative & Ranking Recommender
4:10–4:25 PM Paper Presentation 1: The Disconnect Between Better Descriptive Reasoning Trace Quality and Recommendation Effectiveness — Gustavo Penha
4:25–4:40 PM Paper Presentation 2: Ad Insertion in LLM-Generated Responses — Shengwei Xu
4:40–5:30 PM Poster Session

Keynote Speakers

Patrick Jordan

Patrick Jordan

Microsoft
When Everyone Has an Agent: From Better Decisions to Better Markets

Abstract
Abstract: AI agents are rapidly becoming capable of searching, recommending, negotiating, and transacting on behalf of users and businesses. But something fundamental changes when these agents interact with other independently motivated agents: the problem shifts from making good decisions to making good decisions in a strategic environment. This talk explores that transition through the lens of recommendation games. When a user agent solicits recommendations from multiple competing agents, each recommendation is no longer simply a prediction of what the user wants—it is also a strategic action shaped by what other agents recommend and by the rules governing selection and reward. More broadly, this talk considers a shift from decision theory to game theory and mechanism design for agentic commerce at scale. The central challenge is not simply to build smarter agents. It is to design and evaluate the environments in which they interact—so that useful behavior becomes strategically advantageous.
                                                                                                                                                                                               
Akshay Soni

Akshay Soni

Shopify
Foundation Models for Agentic and Counterfactual Decision Support in E-Commerce

Abstract
Abstract: Most generative recommendation research models the consumer side, with sequences of a buyer’s clicks, views, and purchases. We describe a foundation model for the merchant side of Shopify, built to support agents that plan and act on a shop’s behalf. The modeled entity is a shop and the sequence is its full operating history, spanning high-frequency behavior (such as sessions and checkouts), pre-computed aggregates (such as GMV summaries), and sparse lifecycle events (such as subscription and churn changes). Catalog entities such as products and pages are tokenized into semantic IDs (SIDs) via residual-quantized autoencoding, compressing a high-cardinality space into a compact vocabulary shared across similar entities, which curbs sparsity and cold-start. A Hierarchical Sequential Transduction Unit (HSTU) backbone is trained autoregressively, combining next-token prediction with a multi-horizon future-token-set objective. The shared representation yields general-purpose merchant embeddings that feed a growing set of downstream applications, from recommendation to forecasting, and enable reasoning about the effect of interventions: a hypothetical action is inserted into a shop’s sequence, and the model predicts the events that would follow, simulating the outcome of taking that action, such as adopting a paid-marketing channel. This grounds autonomous decisions in observed behavior and lets actions be evaluated before deployment.
                                                                                                                                                                                               
Shengbo Guo

Shengbo Guo

Meta
LLM Ranking in Facebook Verticals: from Content-Based LLM Ranking to a Unified Generative & Ranking Recommender

Abstract
Abstract: Large language models are transforming recommendation across Facebook’s verticals by bringing deep semantic understanding — of items, of user intent, and of preference — directly into ranking. That understanding is most powerful precisely where it matters most: on sparse, cold-start, and 0→1 surfaces, where reasoning over content and user preferences unlocks relevance that behavioral signal alone cannot reach. This keynote traces how we put that capability into production, following an arc from content-based LLM ranking to a single unified generative-and-ranking recommender, grounded in production systems and live experiments jointly built across verticals: Marketplace Jobs and Facebook Groups. Presented in conjunction with Heng Liu (Meta) and Shubhojeet Sarkar (Meta). Part 1 — Ranking in Marketplace Jobs. We first follow one vertical end to end. Jobs began with an early content-based LLM ranker that validated the thesis on a cold-start surface — scoring openings from their content rather than interaction history and delivering large relevance and engagement lifts where behavioral signal was thin. It then advanced to the Jobs Subtab SID Ranker, hierarchical discrete item representations (an RQ codebook aligned to the LLM via knowledge-augmented continued pretraining) that substantially compress each item by roughly an order of magnitude, with downstream adaptation via multi-task SFT and reinforcement learning. It turned a chronological feed into a ranking-driven surface, substantially improving offline ranking quality over the prior text ranker and driving incremental engaged-DAU gains. Part 2 — LLM ranking and recommendation in Facebook Groups. We then turn to Groups, where the same content-first philosophy meets a different surface. The Forum ranker — the first LLM ranker deployed in Facebook Groups — brought content-based ranking to community recommendation: a compact distilled student, trained with a multi-teacher agreement scheme and serving continuous relevance scores via Token-Probability Normalization (score = P(Yes)/(P(Yes)+P(No))), lifting relevance substantially while surfacing personally meaningful posts for users with sparse history. Building on it, early Generative-Recommender (GR) experiments — currently over the text modality — extend the approach from pointwise ranking toward generative retrieval and next-item prediction for community and content discovery. Part 3 — Toward a unified generative & ranking recommender. Across both verticals the trajectory converges on one question: can a single LLM backbone both rank with high precision and generate candidates with high recall, collapsing the traditional retrieval/ranking split? We describe unified architecture, the SID-based architecture that ties ranking and generation into one model across Jobs and Groups and points toward agentic recommendation. Lessons and outlook. We close with practical lessons from taking LLM rankers from prototype to production — how item and Semantic-ID representations, teacher–student distillation, and the choice between pointwise-ranking and generative training objectives shape relevance, with serving efficiency treated as a co-design constraint rather than the main story. We end on where this is heading: extending today’s text-only generative results toward multimodal item understanding and, ultimately, agentic recommendation flows that plan and act on a user’s behalf — the trajectory that e-commerce and content platforms increasingly share, from ranking what exists to generating what fits, within one unified model.
                                                                                                                                                                                               

Accepted Papers

  • EvoRank: LLM-Guided Evolution of Multi-Objective Learning-to-Rank Pipelines
    Rayhan Patel, Shabaz Patel
    Abstract
    Abstract: We present EvoRank, an open autonomous ranking engineer: an LLM-guided evolutionary loop that discovers complete Learning-to-Rank pipelines (features, models, losses, ensembles) for multi-objective e-commerce search. On the Expedia ICDM 2013 dataset, with relevance, conversion, and revenue as competing objectives, three independent runs each converge within 50 iterations (about ten dollars) on interpretable pipelines that beat an Optuna-tuned LambdaMART on 60k held-out queries, an advantage that persists at full data scale and places in the top 6 percent of the original competition. A first campaign, evolving only training objectives, builds the central design rule: it appeared to work on its selection fold (the small dataset it uses to pick winners) while a transfer audit, re-scoring winners on held-out data, showed the gains were almost entirely fitness noise (the randomness of its own scoring), and neither seeded domain knowledge nor richer diagnostic feedback changed what transferred. The deciding quantity is measurable in advance: search-space headroom relative to fitness noise. We package this as a headroom gate that predicts, before any LLM spend, whether the loop will pay off, and we release the system, the auditing tools, and a catalog of failure modes with their guardrails, so teams can apply the procedure to their own ranking stacks.
    PDF Code
  • SR-Agent: An Experience-Driven Agentic Framework for Post-Ranking Strategy Refinement in E-Commerce Recommendation
    Hanchen Yang, Kaiwen Yang, Junpeng Zhuang, Yang He, Keting Cen, Bochao Liu, Zhongbo Sun, An Liu, Zhongteng Han, Chenyi Lei
    Abstract
    Abstract: User experience is a first-class objective in industrial e-commerce recommender systems (RS). Post-ranking strategies, which govern diversity, similarity, and exposure over a ranked list, are widely deployed in industrial RS for their simplicity and low serving cost. However, as the online recommendation environment evolves continuously, these statically configured strategies gradually become stale, thereby degrading the user experience. Refining them typically relies on manual inspection, diagnosis, and updates, making it slow, costly, and difficult to scale or reuse. Although recent LLM-based agents (e.g., RecUserSim, SimUSER, and Self-EvolveRec) offer promising directions, none of them close the full loop of automated, self-evolving strategy refinement. To bridge this gap, we introduce SR-Agent, which, to the best of our knowledge, is the first agentic framework deployed to refine post-ranking strategies in industrial RS. SR-Agent unifies three components: (i) a UserSim agent that applies inspection skills to surface user-perceived bad cases; (ii) an Analysis agent that consolidates recurring bad cases into structured, reusable diagnoses; and (iii) a constrained Strategy Refinement Harness that maps diagnoses to typed and bounded actions, gated by a four-stage reward pipeline with reversible rollback. Deployed on the Kuaishou e-commerce platform, SR-Agent continuously runs this refinement loop and, in a one-month online A/B test, increases order volume by 0.71%, browsing depth by 0.34%, and clicked-category diversity by 0.48%, while markedly shortening the refinement cycle and lowering operational cost.
    PDF Code
  • MASLOW: Multi-Agent Synthetic Data Generation for Underspecified, Low-Resource E-Commerce Classification
    Ling Jiang, Shaobai Jiang, Ankur Gupta, Adalie Palma, Ali Marjaninejad, Xiaoyu Chu, Prithvi Sen
    Abstract
    Abstract: Training classifiers for e-commerce catalog management is bottlenecked by two simultaneous scarcities: underspecified task specifications (a bare URL or one-sentence description) and few or no labeled examples. We present MASLOW, to our knowledge the first multi-agent synthetic data pipeline that handles the full journey from underspecified inputs to labeled training data without requiring clean class labels, curated seeds, or task-specific corpora. Four specialized agents progressively enrich specifications, identify decision boundaries, generate diverse examples including hard negatives, and filter for quality. We evaluate MASLOW on 1,925 low-resource product-matching tasks and 40 unseen product classification tasks, two orders of magnitude beyond typical synthetic data studies. Ablation reveals that upstream context enrichment has outsized impact: removing the context-enrichment stage causes twice the performance drop (−5.8% F1) of removing any other agent, suggesting practitioners should prioritize input quality over generation architecture. Synthetic data complements human labels (+11.1% F1 on the scarcest tasks) and enables cold-start training from specifications alone. For In-Context Learning, task-aligned feedback (MASLOW+AFB) that regenerates examples receiving uncertain downstream predictions recovers quality that naive feedback loses. At 47× lower cost and 31× lower latency than human annotation, MASLOW processes 1,925 tasks in approximately one hour.
    PDF Code
  • AURA: Agentic Diagnosis and Refinement for Production Recommender Systems at Scale
    SungGeun Kim, Abhinav Narain, Daniel Nemirovsky
    Abstract
    Abstract: How and why does a recommender system fail the users it serves? Oftentimes in recommender systems, practitioners are left to improve their algorithms based on a combination of feedback from stakeholder teams, domain expertise, and insights from data analyses. Yet, the nuances of how and where the recommendations are performing well or poorly for the end users are difficult to discern from aggregate quantitative metrics. Whereas these metrics can provide a high-level and incomplete picture, further granularity into the quality of recommendations and their patterns requires reasoning with domain understanding and objectivity, at scale. In this paper, we contemplate this complex conundrum and describe a method and implementation that utilizes the latest AI agentic advances to provide actionable diagnoses and improvements for the production recommender systems. We present AURA (Agentic Understanding and Refinement of recommender Algorithms), an end-to-end agentic system that performs qualitative evaluation at scale and can then generate improvements to our algorithms at the code level. Specialized agents read production engagement logs, from thousands of sessions to millions, and surface patterns and examples of how the recommender fails real users. The next step uses those diagnoses as well as context about the recommender's own code, data, and training pipeline to propose and implement refinements grounded in that codebase. We report the system design, initial tests on production data from two large consumer platforms at a major media-streaming company, safeguards, operational learnings, and early results toward a self-improving recommender system. Finally, the diagnostic gap AURA closes is not specific to streaming. The architecture is built to transfer and every domain-specific element enters through the configuration layer that already ported it between our two platforms. We map it concretely to e-commerce and online-retail recommendation.
    PDF Code
  • The Disconnect Between Better Descriptive Reasoning Trace Quality and Recommendation Effectiveness
    Gustavo Penha, Juan Elenter, Claudia Hauff, Hugues Bouchard, Paul Bennett, Mounia Lalmas
    Abstract
    Abstract: Recent work has focused on improving explicit natural-language descriptive reasoning traces for generative recommendation. This includes systems that augment semantic ID (SID) prediction with chain-of-thought reasoning. However, because SIDs are opaque learned identifiers rather than natural language, they require costly alignment before an LLM can reason over them. This provides a controlled experimental setting in which both item representation (Title vs. SID) and semantic grounding (minimal vs. extensive SID alignment) can be varied independently. We therefore present the first controlled comparison of descriptive reasoning trace quality across semantic IDs and natural-language titles in a 2 × 2 factorial study on three Amazon product domains using a shared Qwen3-1.7B backbone. We find that introducing explicit descriptive reasoning traces reduces traditional offline recommendation effectiveness under standard SFT and RL training, even though natural language titles produce substantially more grounded and interpretable traces. Extensive SID alignment improves descriptive trace quality but not traditional offline recommendation effectiveness, while a richer reward signal partially recovers performance. Overall, our results show that improving descriptive reasoning trace quality is not, by itself, sufficient to consistently improve traditional offline recommendation effectiveness under the training objectives and evaluation protocols studied here.
    PDF Code
  • Efficient Clustering with Quality Guardrails for LLM-based Recommender Systems at Industry Scale
    Longshaokan Wang, Wai Tsang Keung, Punit Ghodasara, Roman Wang, Ali Dashti, Francesc Moreno-Noguer
    Abstract
    Abstract: LLMs can be prohibitively expensive and slow to run at scale, especially for applications that invoke an LLM per sample over millions of inputs. A natural way to scale such applications is to cluster the inputs, run the LLM only on cluster representatives, and propagate the LLM outputs to other cluster members. However, in this setup, the LLM outputs a cluster member receives are only as good as the member's match to the cluster representative. Off-the-shelf clustering methods optimize an aggregate objective over all samples, targeting only average-case quality without offering per-sample quality guardrails. As a result, cluster members can be assigned to poorly-matched representatives, and the inherited LLM outputs — though appropriate for the representative — may be irrelevant or even unsafe for the member. For example, a parent of a toddler grouped with parents of older children could receive age-inappropriate recommendations. Furthermore, most clustering methods scale poorly to millions of inputs in runtime and memory, limiting their applicability at industry scale. We propose a scalable two-stage clustering algorithm with provable per-sample quality guardrails: every sample is guaranteed to share a user-specified minimal similarity and exact attribute match with its cluster representative. The algorithm first generates initial clusters with Mini-batch K-Means, then greedily selects representatives within each initial cluster to satisfy the guardrails. We provide theoretical guarantees, complexity analysis, and benchmarks against common clustering methods on internal and public datasets. We show that our method not only delivers per-sample guardrails but also runs substantially faster and scales to data sizes where most standard methods become intractable. We demonstrate our method's impact in a real-world deployment clustering 38 million customers. We reduce downstream LLM cost and runtime by 50× while preserving personalization, unblocking the launch of a persona-based recommender system that delivers significant gains in revenue and engagement in an A/B test.
    PDF Code
  • Ad Insertion in LLM-Generated Responses
    Shengwei Xu, Zhaohua Chen, Xiaotie Deng, Zhiyi Huang, Grant Schoenebeck
    Abstract
    Abstract: Sustainable monetization of large language models (LLMs) remains a critical open challenge. Traditional search advertising, which relies on static keywords, fails to capture the fleeting, context-dependent user intent—the specific information, goods, or services a user seeks—embedded in conversational flows. Beyond the standard goal of social welfare maximization, effective LLM advertising requires contextual coherence (aligning ads semantically with transient user intent), computational efficiency (avoiding user-facing latency), and adherence to ethical and regulatory standards, including privacy preservation and explicit ad disclosure. Although recent solutions have explored bidding at the token and query levels, neither category holistically satisfies these constraints. We propose a framework that resolves these tensions through two decoupling strategies. First, we decouple ad insertion from response generation to facilitate pre-screening and explicit disclosure. Second, we decouple bidding from specific user queries by using "genres" (high-level semantic clusters) as a proxy. This allows advertisers to bid on stable categories rather than sensitive real-time responses, reducing computational burden and privacy risks. Applying the VCG auction mechanism to this genre-based framework provides approximate guarantees for dominant-strategy incentive compatibility (DSIC), individual rationality (IR), and social welfare. In synthetic experiments with 105 advertisers and 100 candidate slots, VCG clears in approximately 1.25 seconds on a consumer-grade laptop. Finally, we introduce an "LLM-as-a-Judge" metric for estimating contextual coherence. Its predictions correlate with mean human ratings at Spearman's 𝜌 ≈ 0.66 and have a higher correlation with the leave-one-out group mean than 29 of 36 individual raters (80.6%).
    PDF Code
  • Benchmarking Models for Conversational E-Commerce: A Reproducible Evaluation Framework
    Yuri Vorontsov, Diogo S. Carvalho, Anastasia Vorontsov, Anna Platonova, Vladimir Gorovoy, Ilya Briskin, Felix Tseitlin, Senka Krivic, Salman Ahmad
    Abstract
    Abstract: We present a reproducible framework for benchmarking large language models as multi-turn conversational e-commerce agents across the full shopping loop of search, clarification, recommendation, and add to cart. We demonstrate the framework in a first-in-class 12-model study, complemented by a smaller cross-category evaluation. The framework comprises product catalog ingestion, scenario generation grounded in real consumer search queries, a model-breaking scenario selection filter, LLM-driven customer simulation, and a five-rubric LLM judge covering intent understanding, clarification behaviour, recommendation quality, add-to-cart correctness, and grounding fidelity. We apply the framework to 12 models on a primary jeans case study (1,200 scored conversations across five runs and 20 scenarios), using it as a worked example of the shopping loop; these scores are not intended to generalise across e-commerce. We also conduct a cross-category study on three additional product domains under a different scenario protocol. Rank order on that smaller study is directional only (𝜌 = 0.40–0.70, 𝑛 = 5). On the primary study, grounding fidelity is the main separator between score bands, Clarification is the least saturated rubric among ranks 1–7, and the highest-scoring models form overlapping bootstrap confidence bands once run-to-run variance is accounted for. These findings rest on a judge-reliability programme: a seven-judge cross-family panel with super-judge arbitration that reduces judge swing from 24.5% to 9.2%, validated against a preliminary human gold calibration (103 rubric cells from a single rater; 11 pass / 92 fail). GPT-4.1 alone is too lenient (𝜅 = −0.13); the reliable stack reaches 88.3% accuracy against human labels, below an always-fail dummy at 89.3%, so Cohen's 𝜅 (0.273) is the more informative agreement figure. Benchmark code and data are at https://github.com/rezolved/conversational-commerce-benchmark-framework.
    PDF Code
  • Domain-Adaptive Contrastive Embeddings for Frugal Skill Routing in Agentic E-Commerce Evaluation
    Yue Yu, Meisam Hejazinia, Mark Martinov Kirichev, Siamak Rajabi
    Abstract
    Abstract: Large-language-model (LLM) agents increasingly operate and evaluate e-commerce workflows, and a recurring building block is a skill router: a semantic-retrieval layer that maps a free-form question to the correct skill in a catalog. The same layer lets an LLM-as-judge evaluator decide which skill should have answered and whether the catalog covers the question at all. In practice it runs on general-purpose embedders that mishandle the dense, abbreviation-heavy vocabulary of retail supply-chain operations, and the reflexive fixes, a larger general embedder or a frontier LLM prompted to route, are metered per call. We instead adapt a small open-source encoder to the domain and deliver the adaptation in the two forms a deployed system needs. The static recipe fully fine-tunes a 22M-parameter MiniLM-class encoder on typed contrastive pairs under a leakage-free skill-holdout split, lifting overall routing Hit@1 from 0.474 (a production 1024-d general embedder) to 0.656, and in-distribution Hit@1 from 0.393 to 0.746, over a compact 384-d index at zero marginal inference cost. The continual-learning recipe keeps that encoder current as the catalog grows from 23 to 43 skills: warm-starting plus a 30% replay sample holds overall Hit@1 near 0.76, matching a full retrain at roughly half the compute while avoiding catastrophic forgetting. We state the main caveat plainly: the gains concentrate on skills seen in training, and on genuinely unseen skills the general embedder still leads. A score-fusion ablation tops every arm but reintroduces a per-query general-embedder call, so we keep the standalone specialist central and point to a lightweight learned gate as future work. For the retrieval substrate of agentic e-commerce, a small specialized embedder is both more accurate in domain and far cheaper than a larger general model or a frontier LLM.
    PDF Code
  • What Matters for LLM Persona Simulation: Signal Form in Cold-Start E-Commerce Advertising
    Jiwoo Choi, In Ik Lee, Woo Sung Seo, Minho Lee, Yun Young Choi
    Abstract
    Abstract: Estimating advertising performance in cold-start settings is challenging because historical logs are unavailable before launch. Large language models (LLMs) offer a potential alternative by simulating persona-conditioned user responses, but it remains unclear which design choices actually drive LLM-based persona simulation. Using real Google Ads e-commerce campaign data across five demographic segments, we identify three signal-form dimensions that affect simulation quality: (i) a demographic anchor without noisy rich attributes, (ii) a relative qualitative banner-blindness signal rather than quantitative anchors, and (iii) prompt phrasing. On this campaign, our method attains an MAE of 24.9 and a MAPE of 26.1%, a 51.3% lower MAE than the best tested non-LLM baseline and 77.3% lower than an LLM configuration using explicit CTR anchors. Systematic ablations show that minimal demographic descriptions achieve lower mean error than richer persona configurations, while relative qualitative cues substantially outperform the tested quantitative anchors. We additionally compare these configurations with direct-prior and rule-based baselines and with alternative LLM configurations.
    PDF Code
  • Beyond Ranking Accuracy: Evaluating LLM-Cited Feature Rationales for Next Basket Repurchase Recommendation
    Yanan Cao, Anay Dombe, Murali Mohana Krishna Dandu, Shreeranjani Srirangamsridharan, Sinduja Subramaniam, Yogananth Mahalingam, Evren Korpeoglu, Kannan Achan
    Abstract
    Abstract: Next-basket repurchase recommendation is commonly formulated as a ranking task: given a customer's purchase history, the system ranks previously purchased items that may be needed again. In production settings, however, ranking accuracy is only one component of recommendation quality. Customers may also benefit from concise evidence about why an item is recommended now. Large language models (LLMs) offer a potential mechanism for surfacing such evidence by producing feature-based, human-readable rationales grounded in interpretable behavioral signals rather than ranking scores alone. We construct interpretable repurchase features spanning cadence, frequency, recency, user behavior, and item popularity, and evaluate LLMs on two public grocery datasets and one proprietary retail dataset. We investigate (1) whether off-the-shelf LLMs can use these features as next-basket scorers relative to personal-frequency and supervised rankers, and (2) whether the features cited by LLMs as rationales carry outcome-grounded ranking signal. For the latter, we compare LLM-cited features with model-specific attribution methods under a cross-model feature-masking protocol that measures ranking-quality degradation after masking selected features. Our results provide a nuanced view of LLM scorers for repurchase recommendation. LLM scores are not competitive with supervised rankers, suggesting that off-the-shelf LLMs should not be used directly as standalone repurchase recommenders. However, changes in prompt and evidence representation can yield stronger outcome-grounded feature-masking results in some settings even when ranking performance does not improve; this pattern is dataset-dependent and does not consistently match model-specific attribution baselines. These findings suggest a practical role for LLMs as validated explanation components: not as primary next-basket rankers, but as tools for surfacing behavioral evidence whose outcome-grounded relevance should be evaluated separately from ranking accuracy.
    PDF Code
  • Distill Globally, Adapt Locally: Reasoning Distillation and Product-Type Test-Time Training for Scalable Trade-Up Recommendation
    Siliang Liu, Mohammad Ghasemi, Sapan Patel, Amin Banitalebi-Dehkordi
    Abstract
    Abstract: Trade-up recommendation aims to identify higher-quality alternatives that preserve a customer's purchase intent while offering upgraded benefits through improved formulation, certifications, or brand positioning. Although large language models (LLMs) can reason about these subtle distinctions, applying them directly to hundreds of millions of product pairs in e-commerce stores is operationally impractical. We introduce a two-level framework that distills LLM-derived reasoning into an efficient non-generative student and subsequently adapts its decision boundary to product-type-specific trade-up criteria. At Level 1, a retrieval-augmented few-shot LLM teacher generates both structured relation labels and natural-language rationales. These rationales are encoded and transferred to a compact, non-generative embedding-pair classifier through alignment and contrastive objectives. At inference, the student consumes only two precomputed 768-dimensional product embeddings, requiring neither LLM calls nor text generation. On a fixed human-annotated benchmark of 8,352 pairs, a 15.5M-parameter four-class reasoning-distilled student achieves an AUC of 0.924 (95% CI [0.918, 0.929]), improving over the four-class label-only student with the same shallow architecture (AUC 0.912); rationale supervision provides little benefit when the same task is collapsed to binary labels. At Level 2, we introduce product-type test-time training (PT-TTT), which uses few-shot demonstrations as gradient-based supervision to optimize lightweight category-specific adapters over the frozen student model. PT-TTT improves AUC from 0.924 to 0.941 and average precision from 0.920 to 0.940 without serving-time LLM inference. On a 100K-pair proxy catalog, inference with the distilled student on a single eight-GPU machine is approximately 5,000× faster and has an estimated cost approximately 10,000× lower than direct LLM inference on the same workload.
    PDF Code
  • DENIM: Ablating Catalog Access and Visual Grounding in an Agentic Fashion Shopping Assistant
    Alessandro Francesco Maria Martina, Cataldo Musto, Marco de Gemmis, Pasquale Lops, Giovanni Semeraro
    Abstract
    Abstract: Large Language Models now power e-commerce shopping agents that converse with customers, call retrieval tools, and recommend from live product catalogs. Building one means deciding how the product catalog reaches the agent: what it can query or read at inference time, and in what representation. We isolate two of these decisions: what kind of catalog access the agent needs, and in what form visual item information should reach it. To answer both questions, we first implemented an agentic conversational recommender for fashion e-commerce. It is designed as a single orchestrated system with tool-augmented retrieval over a large fashion catalog. We then ran controlled single-change ablations on it, evaluated with a leakage-controlled simulated shopper in a cold-start setting. First, structured catalog access is the dominant factor: removing it costs 31 points of success rate, and most of that value (27 points) comes from a catalog inspection tool that lets the agent see what the catalog contains and which attribute values are valid before filtering, a capability that recent systems include in various forms but whose effect on retrieval had not been isolated. Second, what matters about visual information is the form in which it reaches the agent: injecting item images at runtime adds no measurable success over offline-extracted visual attributes, while removing those attributes' structured rendering costs a further 11 points even though the same information remains available in prose descriptions. At matched item representation, the orchestrated pipeline also outperforms ReAct and Reflexion-style baselines built on the same backbone and tools by 12–20 points, sustaining preference elicitation where they stall. All findings replicate in direction across two simulator LLM families, one of which shares the recommender's own model family. We distill the results into practical guidance for building e-commerce shopping agents.
    PDF Code

Organizers

Mansi Mane

Mansi Mane
Walmart

Bio
Bio: Mansi Mane is Staff Machine Learning Scientist at Search and Recommendation team in Walmart. She was the main organizer for the first and second workshop for Generative AI for E-commerce. She completed her Masters from Carnegie Mellon University in 2018. She currently focuses on research and development of machine learning algorithms for recommendations, search, marketing as well as content generation. Mansi was previously Applied Scientist at AWS where she led efforts for training of billion scale large language models from scratch. Her research interests include machine learning, multimodal LLMs pretraining, fine-tuning as well as in-context learning. She has published papers in ICML, RecSys, WWW conferences.
Neeti Narayan

Neeti Narayan
Amazon

Bio
Bio: Neeti Narayan is a Senior Applied Scientist at Amazon, leading efforts focusing on GenAI-based commerce content, personalization, and product recommendation systems that have driven millions of dollars in revenue. She also organizes workshops at Amazon’s internal Machine Learning Conference (AMLC) on the application of generative AI in advertising, and actively contributes to broader scientific community through paper reviews and collaborative research. Prior to Amazon, Neeti held a research position at Yahoo, where she implemented and deployed large-scale email classification and user action prediction models processing billions of emails every day. She holds a Ph.D. degree from SUNY Buffalo (2018). Her interests span multimodal LLMs, NLP, and computer vision. Neeti has published in conferences and journals such as CVPR, ISBA, and Image and Vision Computing.
Djordje Gligorijevic

Djordje Gligorijevic
Meta

Bio
Bio: Djordje Gligorijevic is applied sciences manager at Meta, leading Intelligent Harvesting team in Meta’s Ranking AI Monetization organization focused on developing model templates with state-of-the-art ML techniques and agentic model optimization for the entire ads ecosystem. Previously he worked as Applied Research Manager at eBay, and as a Research Scientist in Yahoo Research. He received the Ph.D. degree from Temple University, Philadelphia, PA, in 2018. His research interests include Machine Learning, Extreme Multi-Label Classification, NLP, LLMs, and the Integration of Qualitative Knowledge into predictive models with applications in domains of Computational Healthcare, Computational Advertising, Search, Ranking, and Recommendation Systems. Djordje has published at international conferences such as AAAI, KDD, TheWebConf, SDM, CIKM, SIGIR, as well as in international journals like Data Mining and Knowledge Discovery, BigData journal where he serves as associate editor, Methods and Nature’s Scientific Reports.
Dingxian Wang

Dingxian Wang
Upwork

Bio
Bio: Dingxian is an Applied Science leader with around 12 years industry experience at the intersection of machine learning, software engineering, applied science, and product development. He is passionate about applying skills to solving real-world problems, especially in the field of technology and data science. He is currently leading a team focus on the ranking, personalization and recommendation in the search area. Throughout the career, Dingxian has been involved in a wide range of areas, including search engine, query understanding, recall system, ranking system, recommender system, marketing science, personalization, information extraction, knowledge graph etc. With massive proven track records of delivering great business results, and drove hundreds of millions of dollars in GMV and revenue growth. Dingxian has received many top honors and awards ranging from top conference, journals, patents to top research projects as well as internal competition awards. Including 20+ papers on top conference and journals (one best paper candidate of CIKM 2021), 9 US patents, over 2500 citations, ICT Research Project of the Year 2021 of ACS (Australian Computer Society), Upwork All Star Award and eBay Leaders’ Choice Award.
Topojoy Biswas

Topojoy Biswas
Walmart

Bio
Bio: Topojoy Biswas is Distinguished Data Scientist at Walmart. At Walmart he leads efforts related to W+ membership models and creative generation projects. Prior to Walmart he worked as Principal Engineer at Yahoo Research where he worked on information extraction on text and videos in Yahoo Knowledge Graph which powers search and information organization in products in Yahoo, like Finance, Sports, entity search and browse. Before Yahoo Knowledge graphs, he worked for Yahoo shopping on attribute extraction and classification of shopping feeds into large taxonomies of products. Topojoy has published in multiple international conferences such as ICIP, ACM Multimedia etc and has spoken on applied machine learning topics in MLConf, KGC etc.
Claudio Pomo

Claudio Pomo
Politecnico di Bari

Bio
Bio: Claudio Pomo is an assistant professor at Politecnico di Bari working on responsible AI for personalization, with a focus on reproducibility and evaluation in recommender systems. He has published at venues such as SIGIR, RecSys, ECIR, and UMAP, and in journals including ACM TORS, TKDE, Information Sciences, and IP&M. He organized the “EvalRS 2023” workshop at KDD, chaired the RecSys Challenge in 2024 and 2025, and co-organized the first edition of “DaQuaMRec” workshop at RecSys 2025.

Contact

For any questions, please email genai-ecommerce@googlegroups.com.