FAQ

This page is for short, practical answers.

For the bigger methodological questions, start with:

Installation and setup

How long does the first Stan compilation take?

Usually 1 to 3 minutes. Subsequent runs typically reuse the cached binary. If compilation appears stuck, check the C++ toolchain in Install and Setup.

Do I need to set R_LIBS_USER every time?

Yes, unless you add it to your shell profile. Use dsambayes_set_r_library host to select a repo-local path that is isolated from both your system library and any container library.

Can I use renv instead of .Rlib?

Yes. The repo includes renv.lock. Use renv::restore() if you want exact dependency restoration.

Modelling

How many weeks of data do I need?

There is no hard minimum, but a useful rule of thumb is:

  • BLM: about 100+ weeks for a model with roughly 10 to 15 predictors
  • Hierarchical: about 80+ weeks per group, ideally with at least 4 groups

Shorter series can still be modelled, but the posterior will usually be much more prior-driven and less decision-ready.

Should I use identity or log response?

  • Identity when the KPI is naturally additive and variance is fairly stable
  • Log when the KPI is strictly positive and effect interpretation is more naturally multiplicative

If unsure, fit both and compare the adequacy and diagnostic picture, not just a single fit metric. See Response Scale Semantics.

How should I think about priors?

Start with Stage 2: Model and Priors.

The short answer is:

  • start with defaults
  • add sparse sign constraints only for structural assumptions
  • add sparse prior overrides only when you can defend them in business or modelling terms
  • do not use priors as a substitute for weak model design

Where do prior numbers come from, and why aren’t they standardised like some other MMM tools?

Short answer: in DSAMbayes, you write priors in the original units of your data, for example revenue, not in a standardised or z-score space. This is a deliberate design choice, not a mistake, and it differs from some other MMM tools, including Abacus, which ask you to specify priors directly in their internal standardised space.

What “standardised space” means

Before Stan samples, most Bayesian MMM tools, DSAMbayes included, centre and scale the response and predictors so the sampler works with unit-scale quantities. This is a computational convenience: it improves sampler efficiency and keeps step sizes well behaved across parameters with very different natural magnitudes (e.g. revenue in the hundreds of thousands versus a spend variable in the low thousands).

Where tools differ is which layer of that process they ask the user to think in.

  • Abacus-style: you write the prior directly on the standardised space. A config entry such as intercept: Normal(mu = 0, sigma = 2) is only interpretable because Abacus has already decided how the target is scaled; the 2 is “2 standard deviations of the standardised target”, not “2 units of revenue”.
  • DSAMbayes: you write the prior on the original, reporting-scale units of your data. intercept: Normal(mu = 120000, sigma = 15000) means “baseline revenue is around 120,000, plus or minus roughly 15,000”, in whatever currency or unit your target column uses. When model.scale: true (the default), DSAMbayes converts this raw-scale prior into the correct standardised parameterisation internally, fits, and then back-transforms the posterior draws to your original units automatically. You never see or touch the standardised numbers unless you set model.scale: false.

Both approaches produce equivalent underlying sampler behaviour. Neither is “unscaled” in the sense of skipping standardisation; DSAMbayes just does not require you to author priors in the post-scaling space.

Why DSAMbayes uses the reporting scale for the prior contract

  1. The numbers stay checkable by a non-modeller. A stakeholder, or a colleague reviewing your config, can look at mu = 120000, sigma = 15000 and immediately judge whether that is a sane assumption about baseline revenue. Nobody outside the modelling team can sanity-check Normal(0, 2) on a standardised scale without first recomputing what that implies in revenue terms.
  2. The prior’s meaning does not silently drift between runs. A standardised-space prior such as Normal(0, 2) means something different each time you refit, because the underlying standard deviation used to derive that scale changes with the data (different date range, different market, a new data pull). A reporting-scale prior such as Normal(120000, 15000) means the same thing, “baseline revenue is about 120k”, regardless of what the data’s standard deviation happens to be on any given run. This avoids a common failure mode where a prior that looked reasonable on one dataset becomes accidentally very informative, or accidentally very diffuse, on the next.
  3. It matches how priors are usually elicited in practice. Domain knowledge about a KPI baseline, a channel’s plausible ROI, or a competitor effect is almost always expressed in business units. Asking the analyst to translate that into a standardised space by hand is an extra, error-prone step that this design avoids.

This is consistent with standard Bayesian workflow guidance (see What Principled Means and Stage 2: Model and Priors): elicit priors where you have real, defensible intuition, and let the software handle re-parameterisation for computation.

What actually happens under the hood

DSAMbayes’s internal scaling of priors and boundaries, including the exact transformation ratios for slopes, the intercept, and hierarchical group-level standard deviations, is fully documented in Priors and Boundaries under “Scale semantics”. That page is the technical reference; this FAQ entry is only the “why”.

How to sanity-check a prior you have written

  • Reason about it directly in your data’s natural units; that is what the YAML priors: block and set_prior() both expect.
  • Use peek_prior(model) to see the full prior table on the reporting scale before fitting.
  • If you want to see the internal standardised-space numbers DSAMbayes actually passes to Stan, set model.scale: false on an otherwise identical model and compare, or inspect the scaling terms directly; this is rarely necessary for day-to-day modelling.
  • Do not manually convert your prior into a standardised space and enter that instead. That would be double-scaling and will produce the wrong prior.

See also Minimal-Prior Policy for the recommended default-first, sparse-override operating rule that this scaling behaviour is designed to support.

How many MCMC iterations do I need?

The defaults are a reasonable starting point. Then inspect the Stage 4 diagnostics:

  • Rhat <= 1.01 and healthy ESS usually mean the draw count is adequate
  • Rhat > 1.01 or weak ESS usually means you need to increase iterations and warmup
  • any divergences should be addressed before treating the fit as decision-ready

See Stage 4: Computation and Sampler.

How strict is the stationarity requirement for MMM?

DSAMbayes does not require the raw KPI to satisfy a textbook stationarity condition before fitting.

The important question is whether the remaining unexplained structure, after adding sensible controls and baseline terms, is weak enough that media effects are not standing in for missing baseline dynamics.

When should I set boundaries on media coefficients?

Use m_channel > 0 when non-negativity is a structural belief you would defend in writing. Do not apply blanket sign constraints just to make the output look tidier. See Stage 2: Model and Priors and Minimal-Prior Policy.

When should I use CRE (Mundlak)?

Use CRE when you want to separate within-group temporal effects from between-group cross-sectional structure in a hierarchical model. See CRE / Mundlak.

How should I handle CRE mean terms in decomposition / attribution?

Treat cre_mean_* terms as baseline or between-group structure, not as media attribution terms. They are there to absorb confounding structure, not to claim channel contribution.

What priors should I use on CRE mean terms?

Usually the defaults. Avoid manually tightening or positivity-constraining them unless you have a very strong reason, because that can undermine the whole point of CRE adjustment.

Can I add random slopes for CRE mean terms?

No. Those terms are constant within group, so random slopes on them are not separately identifiable from the group intercept.

What does scale = TRUE do?

It standardises the response and predictors before Stan fitting to improve sampler efficiency. Post-fit coefficient extraction is back-transformed automatically.

Runner and outputs

How long does a typical run take?

Roughly:

  • BLM MCMC: a few minutes
  • BLM MAP: seconds
  • Hierarchical MCMC: tens of minutes depending on size
  • Pooled MCMC: usually between BLM and hierarchical

First-time Stan compilation adds extra startup time.

What is the difference between validate and run?

  • validate checks config and data contracts without compiling or fitting Stan
  • run validates, fits, writes staged artefacts, and runs diagnostics

Always validate first after config changes.

Where do outputs go?

Under results/<timestamp>_<model_name>/ by default. See Output Artefacts.

How do I compare two model runs?

Use compare_runs() or compare the model-selection artefacts directly. See Compare Runs.

Diagnostics

Which diagnostics matter most?

Read them in order:

  1. Stage 4: Computation and Sampler
  2. Stage 5: Model Adequacy

That is more important than memorising one threshold in isolation.

What does “Pareto-k > 0.7” mean?

It means the LOO approximation is unreliable for that observation and the point is highly influential. Investigate the observation and be cautious about using LOO-based comparisons mechanically.

My diagnostics say warn. Should I worry?

Usually yes, but not always in the same way.

  • in exploratory work, a warning may be acceptable if understood
  • in shareable reporting, warnings should be disclosed and interpreted
  • repeated or severe warnings usually mean the model needs revision before decision use

Use Interpret Diagnostics for triage and the workflow pages for meaning.

Budget optimisation

How does the allocator work?

It searches feasible spend allocations within channel constraints and scores them against the fitted model. It is a decision layer built on the model, not an independent source of truth.

Can I use budget optimisation with MAP-fitted models?

Yes, but then the result is point-estimate-driven rather than uncertainty-rich. That is fine for rough iteration, not ideal for final decision support.