Abstract

Many empirical applications observe outcomes as event logs: transactions, calls for service, incidents, clicks, visits, or other timestamped arrivals. The usual difference-in-differences workflow bins these events into daily, weekly, or monthly counts and then estimates a linear two-way fixed-effects regression. The choice of time bin is inseparable from the scale on which the identifying restriction is imposed. For event data, a useful baseline model is a counting process with untreated intensities that factor into unit and calendar-time components. Under a proportional treatment effect on the intensity, the causal estimand is a rate ratio. A log rate-ratio difference-in-differences estimator and a Poisson fixed-effects estimator recover this object under multiplicative parallel trends. Linear fixed effects on binned counts instead impose additive parallel trends on the count scale; in general it estimates an additive rate or count contrast, not the log rate ratio. Changing the bin width cannot repair this scale mismatch and, for raw counts, mechanically changes the magnitude of the OLS coefficient. The analysis connects this point to Poisson pseudo-maximum likelihood in econometrics and to recurrent-event methods in biostatistics, and gives practical recommendations for event-log DiD designs.

Introduction

Event logs are now a common form of economic data. A platform records user actions; a city records calls for service; a hospital records admissions and readmissions; a firm records transactions; a website records visits. In all of these examples the primitive object is not a scalar outcome observed once per period but a timestamped point process. Applied work usually converts that process into a panel by choosing a calendar interval and counting events in each cell. The resulting outcome is then often analyzed with the same linear two-way fixed-effects tools used for continuous panel outcomes.

The conversion from event times to panel counts changes more than the sampling frequency. It also fixes the scale of the identifying restriction. When untreated event intensities follow multiplicative parallel trends, differences in baseline event rates across units are proportional rather than additive. If treatment has a proportional effect on the event intensity, the corresponding treatment effect is a log rate ratio. Poisson pseudo-maximum likelihood and a simple log rate-ratio DiD estimator are built for this scale. Linear regressions on binned counts are not.

The argument has four steps. First, a potential-outcome framework for counting processes leads directly to rate-ratio estimands for event data. Second, under untreated multiplicative parallel trends, a four-cell log rate-ratio estimator identifies the proportional treatment effect even when treated units have different baseline rates. Third, Poisson fixed effects consistently estimate the same log rate ratio when the binned conditional mean is exponential in unit, time, and treatment effects. The claim is quasi-likelihood: the conditional distribution need not literally be Poisson. Fourth, OLS on counts answers a different question. It can be a reasonable estimator of an additive count effect if additive parallel trends on counts is the target assumption, but it should not be interpreted as a percent lift or log rate ratio merely because the original data are finely binned.

Before binning event data, researchers should state whether the estimand is an additive change in expected counts, an additive change in rates, or a multiplicative change in rates. For most “lift” questions, the estimand is multiplicative. In those cases log rate-ratio DiD or Poisson/PPML fixed effects are the appropriate estimators, with exposure handled explicitly and with bins chosen so treatment status is not ambiguous within a cell.

Figure 1 gives the basic two-period logic. The treated group has a higher baseline event intensity than the control group, both groups experience a common inhomogeneous time pattern, and treatment raises the treated group’s post- intensity by 5 percent. The vertical gray lines are possible bins. The four numbers in the DiD contrasts are group-period average event rates: for control pre, for control post, for treated pre, and for treated post. With unit-width bins these are expected counts per bin; more generally they are expected counts divided by exposure. The log rate-ratio DiD recovers the 5 percent treatment effect because it compares proportional changes in rates, not differences in raw counts. The vanilla levels DiD is also well defined, but it reports rate units; normalized by the control pre-period mean, this is percent, far from the true 5 percent proportional lift.

Figure 1. Event-time view of a two-period rate-ratio DiD. Dots are event times for treated and control groups. Gray lines show possible time bins; the dotted line marks treatment. The four numbers in the contrasts are group-period average event rates. The two-period log rate-ratio DiD equals the true 5 percent proportional effect; the levels DiD reports a different additive rate contrast, equal to 32 percent when normalized by the control pre-period rate.

Figure 1. Event-time view of a two-period rate-ratio DiD. Dots are event times for treated and control groups. Gray lines show possible time bins; the dotted line marks treatment. The four numbers in the contrasts are group-period average event rates. The two-period log rate-ratio DiD equals the true 5 percent proportional effect; the levels DiD reports a different additive rate contrast, equal to 32 percent when normalized by the control pre-period rate.

The econometric foundation for this work is the pseudo-maximum likelihood literature. Poisson pseudo-likelihood can deliver consistent and asymptotically normal estimators under conditional-mean restrictions even when the full distribution is misspecified (Gourieroux, Monfort, and Trognon 1984a,b). Panel count models entered applied econometrics early in the patents and R&D literature (Hausman, Hall, and Griliches 1984). Fixed-effects Poisson estimators have useful robustness properties in multiplicative panel models when the conditional mean is correctly specified (Wooldridge 1999). The same logic became central in the trade literature: when the conditional mean is multiplicative and heteroskedasticity is present, log-linear OLS is generally not estimating the desired conditional-mean parameter (Santos Silva and Tenreyro 2006). Modern algorithms make Poisson models with high-dimensional fixed effects practical at large scale (Correia, Guimaraes, and Zylkin 2020). Recent nonlinear DiD work also develops exponential conditional-mean models for panel data (Wooldridge 2023). For timestamped event logs, the same conditional-mean logic applies after binning: exposure length and treatment timing determine the mean structure that PPML estimates.

The biostatistics and event-history literatures provide the stochastic-process language for the primitive outcome. The Cox model has a counting-process formulation, and there is a mature monograph literature on statistical models based on counting processes (Andersen and Gill 1982; Andersen et al. 1993). Recurrent-event methods are well developed (Cook and Lawless 2007), including semiparametric regression for recurrent-event mean and rate functions without requiring a Poisson process assumption (Lin et al. 2000). Competing risks and semicompeting risks have their own estimands, including cause-specific hazards, cumulative incidence functions, and subdistribution hazards; the subdistribution hazard model is the standard example (Fine and Gray 1999). Recent causal work explicitly formulates recurrent-event and competing-event estimands with potential outcomes and counting-process limits; Janvin et al. (2024) is closely related in its use of potential outcomes for recurrent event processes. Much of the causal recurrent-event literature identifies effects under observed covariate adjustment, sequential exchangeability, or structural hazard and mean-model restrictions. The DiD setting considered here instead allows treatment status to be related to time-invariant latent baseline intensity, provided untreated intensities satisfy a multiplicative parallel-trends restriction.

Relative to these literatures, the contribution is not a new event-history estimand or a new Poisson estimator. Familiar rate and mean-function estimands can be linked directly to DiD practice with timestamped event logs. The identifying restriction is a DiD restriction on untreated potential intensities: time-invariant baseline rate differences may be arbitrarily related to treatment status, but they enter multiplicatively and are removed by a rate-ratio contrast or by fixed effects in an exponential conditional mean. That link clarifies the role of temporal aggregation and explains why OLS on binned counts can fail when researchers intend to estimate a percent effect.

Counting-Process Setup

There are units observed over a fixed horizon . For unit , denotes the number of events in , and . In the main two-period design, some units are treated beginning at and others are never treated. Let denote membership in the treated group and let .

For treatment regime , let be the potential counting process and let be its predictable intensity. The compensator is

so that is a martingale under the usual event-history regularity conditions. The observed process satisfies consistency:

The post-treatment causal estimand is the proportional effect on expected event arrivals among the treated group:

The parameter is a causal rate ratio for cumulative event arrivals over a specified post-treatment window. If treatment has a constant proportional effect on the intensity, then equals that log intensity ratio.

The target is a marginal rate ratio over a window, not a hazard ratio conditional on survival or on the full event history. For recurring events without terminal competing events, this is close to the marginal mean/rate function estimands used in recurrent-event analysis. If terminal or competing events are present, the estimand must be modified because time at risk is itself affected by treatment. That case requires the recurrent-event and competing-risk estimands studied in the biostatistics literature.

Identification in a Two-Period Design

Let and . Write

The identifying restriction is multiplicative parallel trends for untreated rates. There exist group-specific baseline factors and period-specific rate factors such that

This condition allows treated and control units to have different baseline event rates. It restricts the untreated evolution of rates to be proportional across groups.

The treatment effect restriction is proportionality in the treated post period:

Together, these assumptions imply

where denotes the observed group-period rate. The corresponding sample estimator is

with equal to total events divided by total exposure in cell .

The same contrast obtains from a saturated Poisson model for the four cell totals with log exposure as an offset. The identifying assumption concerns expected rates: untreated expected rates are separable in group and period, and treatment multiplies the treated post rate.

Temporal Aggregation and Binned Counts

Applied researchers rarely estimate from the raw event times. They choose a partition of and construct

Let denote exposure in bin . If the untreated intensity has the form

and treatment is constant within bin , then

Equivalently,

Therefore a Poisson fixed-effects regression with unit and bin effects is correctly specified at the level of the conditional mean, even if the baseline intensity varies within the bin. The bin fixed effect absorbs the integrated baseline intensity.

This exact representation requires treatment status to be constant within the bin. If a bin straddles a treatment adoption time, then the exact mean is

which is generally not equal to for a fractional exposure variable . Event-log applications should therefore split bins at treatment adoption times or include separate treated and untreated exposure components when treatment changes inside a bin.

The Poisson estimator solves the sample analogue of the conditional-mean moment condition

Interpreted this way, the estimator is PPML/QMLE. It does not require equidispersion, independent event arrivals, or a literal Poisson process, provided the conditional mean is correct and inference accounts for clustering or serial dependence.

What Linear Fixed Effects Estimates

OLS on binned counts imposes an additive conditional mean:

The linear specification corresponds to additive parallel trends in counts, not multiplicative parallel trends in rates. Except in degenerate cases, additive and multiplicative parallel trends cannot both be true at the same time.

The distinction is visible in the two-by-two population coefficient. With equal exposure within period and untreated means , the saturated linear DiD contrast is

If , this equals , which is generally nonzero whenever treated and control units have different baseline rates and the common rate factor changes over time. If , then

an additive rate or count effect for the treated group, not the log rate ratio .

As bins shrink, a raw count coefficient changes with the exposure length of the bin. If counts are divided by exposure, the limiting object is an additive rate contrast. Neither object is a log rate ratio unless additional approximations are imposed. OLS can be appropriate for an additive count estimand. Without a model or transformation targeting a multiplicative estimand, it is not a percent-effect estimator.

Simulation Evidence

The simulation design varies treatment assignment while holding fixed a multiplicative event-rate data-generating process. Each replication has units observed over periods, with treatment beginning at for treated units. Unit heterogeneity is and the auxiliary covariate is . Conditional on these quantities, events are generated by an inhomogeneous Poisson process with intensity

with a true log rate ratio . The calendar component makes the process inhomogeneous but common across units.

There are three assignment designs. In the first, exactly half the units are treated at random and . In the second, treatment depends on the unit fixed effect:

This creates large baseline rate differences between treated and control units, but those differences are time invariant and therefore compatible with multiplicative parallel trends. In the third, treatment depends on an unobserved covariate:

Here affects both treatment and the baseline event rate, but it is time invariant and enters the intensity multiplicatively, so the maintained DiD restriction is still satisfied.

The estimators are a four-cell log rate-ratio DiD estimator, Poisson fixed effects with fine and coarse bins, and linear fixed effects on raw counts with fine and coarse bins. Because the OLS coefficient is an additive count contrast rather than a log rate ratio, Figure 2 plots the linear estimates on their own scale.

Figure 2. Monte Carlo estimates under three assignment designs. The top row shows estimators whose target is the log rate ratio. The dashed vertical line is the true value \tau=0.5. The bottom row shows OLS fixed-effect coefficients on raw binned counts; these are additive count effects, not log rate ratios.

Figure 2. Monte Carlo estimates under three assignment designs. The top row shows estimators whose target is the log rate ratio. The dashed vertical line is the true value . The bottom row shows OLS fixed-effect coefficients on raw binned counts; these are additive count effects, not log rate ratios.

Across all three assignment designs, the four-cell log rate-ratio DiD and the fine-bin Poisson fixed-effects estimator are centered at approximately . The coarse-bin Poisson estimator is also close, with small deviations caused by coarse bins that can average over changes in treatment exposure or calendar intensity. Selection on and selection on do not move these estimators because both variables shift baseline rates rather than untreated rate growth.

The OLS coefficients are on the raw-count scale. The fine-bin OLS coefficient is an additive count effect per fine interval, and the coarse-bin OLS coefficient is much larger because the outcome is a count over a longer exposure interval. Interpreting either coefficient as a log rate ratio would conflate additive and multiplicative estimands. The design would need additive parallel trends in expected counts, rather than multiplicative parallel trends in event rates, for those coefficients to be the primary causal object.

Event-Log Metrics and xAU Outcomes

Many applications do not use raw event counts. They first transform events into active-user metrics, such as DAU, WAU, or MAU:

Active-user outcomes are another form of temporal aggregation. For a Poisson process with integrated intensity ,

If treatment multiplies the underlying event intensity by , the induced effect on the active-user probability is

which depends on the baseline integrated intensity and the window length. For rare events or very short windows this ratio is close to . For common events or long windows it is attenuated toward one because the active indicator saturates.

Thus xAU effects are not merely rate-ratio effects with a different label. The window length is part of the estimand. Researchers analyzing active-user outcomes should state whether the target is the effect on event intensity, the effect on the probability of at least one event in a specified window, or the effect on the count of active units.

Practical Recommendations

For event-log DiD designs, the following workflow is defensible.

  1. State the primitive outcome: raw event counts, rates per exposure, active-user indicators, first-event times, or recurrent events with terminal competing events.

  2. State the causal scale: additive count difference, additive rate difference, log rate ratio, risk difference, or risk ratio.

  3. If the target is a log rate ratio, use a log rate-ratio DiD estimator in the two-by-two case or a Poisson/PPML fixed-effects estimator in richer panel designs.

  4. Include exposure explicitly. With unequal bin widths or varying time at risk, include log exposure as an offset or use bin fixed effects that absorb the relevant integrated exposure only when appropriate.

  5. Split bins at treatment adoption times. Fractional treatment indicators inside nonlinear mean models are approximations unless the exposure model is written exactly.

  6. Use robust or clustered inference. The Poisson likelihood is a convenient estimating criterion, not a maintained assumption that events are independent and equidispersed.

  7. Treat OLS on counts as estimating an additive count or rate contrast. Do not interpret the coefficient as a percent effect unless the model has been transformed or rescaled to target that object.

Conclusion

Event-log data make the time scale of analysis explicit. The choice of bin width is not a harmless preprocessing step because it interacts with the scale of the identifying restriction and the scale of the estimand. Under multiplicative parallel trends for untreated event rates, the natural DiD estimand is a log rate ratio. Poisson fixed effects and the four-cell log rate-ratio estimator target that object. Linear fixed effects on binned counts generally target an additive count or rate contrast instead. The resulting discrepancy is not fixed by making bins smaller; it is a mismatch between the causal question and the regression mean function.

References

  • Andersen, Per K., Ornulf Borgan, Richard D. Gill, and Niels Keiding. 1993. Statistical Models Based on Counting Processes. Springer.
  • Andersen, Per K., and Richard D. Gill. 1982. “Cox’s Regression Model for Counting Processes: A Large Sample Study.” The Annals of Statistics 10(4): 1100-1120.
  • Cook, Richard J., and Jerald F. Lawless. 2007. The Statistical Analysis of Recurrent Events. Springer.
  • Correia, Sergio, Paulo Guimaraes, and Thomas Zylkin. 2020. “Fast Poisson Estimation with High-Dimensional Fixed Effects.” The Stata Journal 20(1): 95-115.
  • Fine, Jason P., and Robert J. Gray. 1999. “A Proportional Hazards Model for the Subdistribution of a Competing Risk.” Journal of the American Statistical Association 94(446): 496-509.
  • Gourieroux, Christian, Alain Monfort, and Alain Trognon. 1984a. “Pseudo Maximum Likelihood Methods: Applications to Poisson Models.” Econometrica 52(3): 701-720.
  • Gourieroux, Christian, Alain Monfort, and Alain Trognon. 1984b. “Pseudo Maximum Likelihood Methods: Theory.” Econometrica 52(3): 681-700.
  • Hausman, Jerry A., Bronwyn H. Hall, and Zvi Griliches. 1984. “Econometric Models for Count Data with an Application to the Patents-R&D Relationship.” Econometrica 52(4): 909-938.
  • Janvin, Matias, Jessica G. Young, Pal C. Ryalen, and Mats J. Stensrud. 2024. “Causal Inference with Recurrent and Competing Events.” Lifetime Data Analysis 30: 59-118.
  • Lin, D. Y., L. J. Wei, I. Yang, and Z. Ying. 2000. “Semiparametric Regression for the Mean and Rate Functions of Recurrent Events.” Journal of the Royal Statistical Society: Series B 62(4): 711-730.
  • Santos Silva, J. M. C., and Silvana Tenreyro. 2006. “The Log of Gravity.” The Review of Economics and Statistics 88(4): 641-658.
  • Wooldridge, Jeffrey M. 1999. “Distribution-Free Estimation of Some Nonlinear Panel Data Models.” Journal of Econometrics 90(1): 77-97.
  • Wooldridge, Jeffrey M. 2023. “Simple Approaches to Nonlinear Difference-in-Differences with Panel Data.” The Econometrics Journal 26(3): C31-C66.