Onsite Agenda

Auditori de Mercè Rodoreda

Ramon Trias Fargas 25-27

 

 

Thursday 22 Oct 2026

At the School's cafeteria. When you check in at the School's reception desk, you will receive a lunch voucher.

Load your tray with food, hand your voucher to the cashier, and find us in the Fine Dining section.

Ronald Klingebiel, Frankfurt School
Franziska Lauenstein, KLU Hamburg

Though the field of strategy and organization has seen plenty of experimental papers, few go beyond individuals as the level of analysis. Those that do tend to adopt a search (multi-armed bandit) or teams (hidden profile) paradigm, for example. Variation is large. We encourage exploration of stimuli and suggest that it might be too early to settle on workhorses. We do, however, strongly encourage greater borrowing from neighboring disciplines when it comes to basic handicraft issues. Elicitation, incentivization, and participant priors are quick-win areas of improvement.

Room xxx

Hagay Volvovsky, Tel Aviv University

The experiment induces or inhibits shared cognitive schema formation, observing emergence and evaluating the performance consequences when the environment changes. Across two studies, we show that groups that develop shared schemas coordinate faster when environmental change is compatible with those schemas. But when change renderes these schemas ineffective, they impair adaptation. Additional evidence suggests that this maladaptive persistence is strongest when the changed environment remains familiar-looking enough to make an obsolete schema seem useful, and weakest when the schema’s inapplicability is readily apparent. The findings suggest that shared cognitive schemas can limit organizational adaptation and introduce a paradigm for studying how and when schemas become competency traps.

Pia Pustelnik, WU Vienna

We conduct a vignette-based laboratory experiment in which participants assume the role of headquarters managers evaluating investment proposals submitted by foreign subsidiaries. The proposals vary in their level of innovation, ranging from familiar and market-proven initiatives to highly novel opportunities. Participants are randomly assigned to either a human-only decision-making condition or a condition in which they receive AI-generated evaluation support in the form of structured LLM memos. Across both conditions, participants evaluate subsidiary proposals on perceived risk, expected success, strategic fit, innovativeness, and willingness to invest.

Marlon Alves, Skema Business School

Thinking about an experiment on generalists vs specialists, building on a longitudinal study of Stack Overflow.

Hui Sun, Frankfurt School of Finance and Management

We theorize that nomination bias arises both from biased evaluation of the considered candidates (evaluation bias) and from who is considered in the first place (recall bias). Through survey experiments and cognitive modeling, we find that evaluation bias accounts for only a small share of nomination bias; a much larger share arises from recall bias. Further analysis suggests that recall bias is rooted in network proximity rather than socio-demographic similarity, and is immune to changes in nomination tasks. These findings challenge prevailing diversity guidelines on nomination procedures. They suggest that effective efforts should focus less on diversifying nominators per se and more on diversifying their networks and implementing interventions that encourage nominators to consider beyond their close networks.

Mads Kock Skøtt, Aarhus University

Study 1 has participants exposed to payoff landscapes that vary in peak structure and local payoff variability, then complete a prediction task in which they estimate unobserved outcomes given a single revealed solution. We hypothesize that low-variability payoff regions will induce broader generalization than high-variability payoff regions. In Study 2a, participants additionally complete a 1D search task in which they choose solutions and get their fitness value revealed over repeated rounds. Our hypothesis is U-shaped: participants primed toward intermediate generalization breadth search more locally than participants primed toward narrow or broad generalization. In Study 2a, each participant generates a sparse sample of observed locations and payoffs. Because each search path in 2a consists of a limited number of observations, we add human judgments as a measurement. New participants will see observed outcomes from Study 2a and match them to candidate representations that vary in generalization breadth.

Mentees please seek out their mentors as scheduled. Mentors can make it easy by milling around the room.


Please make your own way to

ZZZZ

 

 

Friday 23 Oct 2026

Uri Simonsohn, Esade

Fewer confounds by design: Uri discusses ongoing work on improving the internal validity and credibility of experiments. One stream focuses on stimulus sampling, the idea that experiments should use multiple stimuli per condition. I will discuss how to sample these stimuli systematically and analyze the resulting data informatively. The second stream calls for collecting participants’ explanations in all experiments, to identify unexpected confounds and, sometimes, confirm hypothesized mechanisms. Uri proposes concrete ways to elicit and analyze participant explanations, sharing examples in which the conclusions of published studies are revisited using these newly proposed tools.

Stephan Billinger, Southern Denmark University

The class experiment uses GoVentureCEO, an online business simulation in which students manage firms in competitive product markets. Approximately 550–600 students will complete three five-day simulation runs: first individually, then in three-person teams, and finally individually again. During the team-based run, students will be randomly assigned to one of two conditions. In the fixed-allocation condition, the lecturer assigns responsibility for production, research and development, and sales and marketing to different team members. In the self-selected condition, teams receive the same task structure but decide themselves how to allocate responsibilities.

Almasa Sarabi, Amsterdam University

Thinking about causality to probe results established in a study of job postings at a global tech firm. Findings suggest that inclusive gender cues are positively associated with larger applicant pools, qualified hires, and thereby contributes to a higher likelihood of filling positions.

Thorsten Whale, Skema Business School

We test whether evaluators shift from assessing proposals on evidential merits to assessing the source that generated them, when performance outcomes along a search trajectory are unobservable. This source dependence may create compound distrust (epistemic and motivational) that AI involvement and incentive alignment reduce by mitigating motivational uncertainty.

Jelena Cerar, WU Vienna

We review how emerging technologies reshape experimental possibilities in IB research, highlighting both their methodological advantages and their limitations. Extant lab studies (137) do not yet feature much AI/LLMs, eye-tracking, neuroimaging (fMRI, fNIRS), and virtual reality.

Room xxx

Crowdsourcing best practice in experimentation

Felipe Saez, Manchester University

Xavier Sobrepere, ESCP Madrid

Hui Sun, Frankfurt School of Finance and Management

Hagay Volvovsky, Tel Aviv University

We survey the types of stimuli used for multi-agent lab experiments with organizationally relevant research questions, involving hidden profiles, trust games, or offline negotiation, for example. We discuss the extent of consensus about workhorse stimuli. We report trends in stimulus design for organizational experiments in the lab/online such as a potential bifurcation into incentive compatible studies and those using hypotheticals and priming. We show how stimuli incorporate goal-directedness, joint incentives, organizational structures, decision aggregation, or rivalrous payoffs, for example.

Jelena Cerar, WU Vienna

Franziska Lauenstein, KLU Hamburg

Pia Pustelnik, WU Vienna

We provide a rough measure of the prevalence of organizational stimuli based on hypotheticals, categorizing them. We walk through the implications of using vignettes, including when this is seen as more justified, and when less so. There are specific areas that deserve attention such as task incentives and participant priors.

Mehdi Ibn Brahim, KU Leuven

Almasa Sarabi, Amsterdam University

Mads Kock Skøtt, Aarhus University

To which extent and form is experimental organization science integrated in PhD curricula? Either in business-school syllabi or psych/econ subjects that integrate multi-person experiments in the subjects they teach. Which topics feature? Where would a student need to study to learn the most? What are gaps and low-hanging fruit for extending curricula?

Cafeteria

Rosemarie Nagel, Pompeu Fabra University

Strategic decisions require firms to anticipate the behavior and reasoning of others. Yet laboratory Beauty Contest experiments show substantial heterogeneity in how deeply people reason strategically, giving rise to level- and related models of bounded strategic reasoning. Using experimental evidence as a starting point, how far do these ideas travel from the laboratory to the field? Evidence from managerial and firm decisions — including technology adoption, market entry, bidding, pricing, and information acquisition — suggests that strategic sophistication also varies systematically outside the laboratory and can have important consequences for firm behavior and market outcomes. What can laboratory experiments reveal that field data cannot, what can field evidence teach us about the external relevance of experimental measures, and how to combine both approaches to contribute to a behavioral theory of strategic decision-making in firms?

Room xxx

Crowdsourcing best practice in experimentation

Jerry Guo, Frankfurt School of Finance and Management

Matthias Trobinger, Esade

Thorsten Whale, Skema Business School

We report on any emerging field consensus on (or criticism of) popular subject pools. Frequently sampled pools include those of university labs, business-expert populations, captive class audiences, as well as gig workers [Sona, Sojump, MTurk, Qualtrics, Positly, Prolific, for example]. What is the pools' relative suitability for organizational experiments?

Marlon Alves, Skema Business School

Taeho Kwak, ESSEC

Julian Berger, Pompeu Fabra

We cover the lived norms for tracing causality in organizational experiments. Randomization helps argue that an effect exists but says nothing about why. What is a (sufficiently) clean design revealing causal mechanisms? When and how to measure mediators, when to manipulate them, through moderators that toggle its operation? How to guard against infinite regress? How to choose the right level of abstraction for studying multiply determined organizational phenomena?

Stephan Billinger, Southern Denmark University

Xavier Thuillard, Lugano University

Oana Vuculescu, Aarhus University

Yiheng Zhang, WU Vienna

We round up instances of exploratory experimentation in the broader field of organization science. Given the rich settings of organizational experiments, staging meaningful interventions without clear hypotheses may be appropriate. What are the gold standards for documenting systematic patterns with such an exploratory approach? We examine field norms for choosing which experimental runs to report, how to determine the right number of context-varying sub-studies, among other aspects.

Mentees please seek out their mentors as scheduled. Mentors can make it easy by milling around the room.


Please make your own way to

ZZZZ

 

 

Saturday 24 Oct 2026

Henrik Olsson, Complexity Science Hub Vienna

Belief dynamics in the lab. Abstract tbc

Matthias Trobinger, Esade

Problem-solving through online deliberation is often hindered by premature consensus or deadlock. Theorizing that coordination and communication difficulties are at the root of these challenges, we propose a minimally invasive and decentralized approach: A Shared Attribute Space (SAS) to mentally represent solutions. Across two online experiments involving 788 participants, we find that SAS fosters problem clarity and increases the impact of solutions. However, it also hampers convergence through a contestation dynamic. Our findings contribute to the collective intelligence literature by demonstrating how distributed problem-solving can be made effective while preserving its decentralized ethos.

Oana Vuculescu, Aarhus University

Thinking about a test of anxiety biasing strategic judgment by inducing the misattribution of epistemic uncertainty (resolvable knowledge gaps) to aleatory uncertainty (irreducible randomness). The central claim is that anxiety produces an "overly volatile representation" of the decision environment, leading executives to interpret probabilistic negative feedback as fundamental shifts in the rules of the game rather than noise. This mechanism resolves an apparent paradox in the literature, where anxiety has been shown to both increase exploration.

Yiheng Zhang, WU Vienna

Junior employees gain confidence in Gen AI answers without sufficient experience to detect subtle errors or context-specific nuances, while senior employees lose the standing to push back without appearing defensive. This could result in a new form of intra-team conflict, "competence disagreement", in which both sides assume they are acting rationally given their own expertise, yet arrive at incompatible beliefs about who actually knows what. We measure competence disagreement and its downstream effects on decision quality using a field survey and a lab-in-the-field experiment with managers.

Xavier Thuillard, Lugano University

We compare how varying degrees of information, ranging from perfect information about a task environment to complete ignorance, influence human decision-making in a strategic search task. Participants are randomly assigned to information conditions between subjects, while task complexity is varied within subjects.

Felipe Saez, Manchester University

We conduct a 2×2 between-subjects laboratory experiment (N = 240) that randomly assigns participants to GenAI assistance during (1) idea generation only, (2) implementation only, (3) both stages, or (4) neither. Working individually via Qualtrics with an embedded multi-turn ChatGPT interface and no task time limits, participants first brainstorm creative ideas addressing a food-waste sustainability challenge; the second stage requires them to develop one of their ideas into a strategic implementation plan.

Room S3.01

tbc

Some say we already have, others say we do not quite yet have, AlphaFold for the mind. Recent attempts with Centaur and others and suggests ways of running experiments on AI instead of humans. It may be the only viable path, given that Claude Code & co can create and run experiments on Prolific within minutes, thus anyways diminishing the chance of data being supplied by actual humans (unless we go the route of pen and paper in the lab).

Cafeteria