AIC and BIC disagreeing unanimously in opposite directions
Posted: 03 Sep 2026, 17:20
I am estimating mixed logit models on a discrete choice experiment and have reached a specification decision where the two information criteria disagree completely, and I would be grateful for views on how others handle this.
The design: respondents choose between two improvement packages or a status quo option. Three non-price attributes, each at three levels, plus a price attribute. Six choice tasks per respondent. The sample is split into six subgroups of roughly 370 respondents each, modelled separately, since the study exists to compare valuations across them.
Both specifications are estimated in preference space at 500 MLHS draws, with a lognormal price coefficient, fixed alternative-specific constants, and one attribute (community engagement) held non-random. They differ only in whether the two levels of a second attribute (seating and comfort) are random:
Model A, 14 parameters: seating levels random normal, repair levels random normal
Model B, 12 parameters: seating levels fixed, repair levels random normal
Model A has the lower AIC in all six subgroups, by 4.3 to 8.8 points. Model B has the lower BIC in all six, by 2.7 to 7.1 points. Every subgroup, both criteria, no exceptions in either direction.
Subgroup AIC A AIC B BIC A BIC B
1 3305.73 3311.05 3385.79 3379.67
2 3532.88 3539.47 3612.56 3607.77
3 3440.82 3449.62 3520.96 3518.31
4 3555.16 3559.48 3635.22 3628.10
5 3256.09 3262.45 3335.35 3330.38
6 3641.82 3649.11 3722.10 3717.92
I understand why this happens arithmetically. At roughly 2,200 observations per subgroup, BIC charges about 15 points for the two extra parameters and AIC about 4. The fit loss from fixing the seating standard deviations is around 7 to 12 points, which falls between the two penalties, so each criterion answers as its penalty dictates. What I am unsure about is what to do with that.
What else I have looked at. In Model A the standard deviation on seating level 2 is significant at the 1% level in all six subgroups, with point estimates between 0.64 and 0.84, so the heterogeneity being removed is not marginal. Model B also required non-zero starting values for the constants in one subgroup, where Model A converged from zeros throughout. Both considerations point toward keeping the parameters, but neither is decisive, and I am conscious of arguing my way toward the answer I happen to prefer.
My questions:
Which model should I do and based on what criteria?
Is there a settled convention on which criterion to prefer in choice modelling when the two disagree, or is it genuinely a matter for judgement in each case?
Where the disagreement is driven purely by the penalty on a small number of parameters, as here, does anyone use a formal test — a likelihood ratio test on the restriction, for instance — rather than relying on the criteria?
Does the significance of the standard deviations being removed carry weight in practice, or is that double-counting evidence the criteria have already used?
The design: respondents choose between two improvement packages or a status quo option. Three non-price attributes, each at three levels, plus a price attribute. Six choice tasks per respondent. The sample is split into six subgroups of roughly 370 respondents each, modelled separately, since the study exists to compare valuations across them.
Both specifications are estimated in preference space at 500 MLHS draws, with a lognormal price coefficient, fixed alternative-specific constants, and one attribute (community engagement) held non-random. They differ only in whether the two levels of a second attribute (seating and comfort) are random:
Model A, 14 parameters: seating levels random normal, repair levels random normal
Model B, 12 parameters: seating levels fixed, repair levels random normal
Model A has the lower AIC in all six subgroups, by 4.3 to 8.8 points. Model B has the lower BIC in all six, by 2.7 to 7.1 points. Every subgroup, both criteria, no exceptions in either direction.
Subgroup AIC A AIC B BIC A BIC B
1 3305.73 3311.05 3385.79 3379.67
2 3532.88 3539.47 3612.56 3607.77
3 3440.82 3449.62 3520.96 3518.31
4 3555.16 3559.48 3635.22 3628.10
5 3256.09 3262.45 3335.35 3330.38
6 3641.82 3649.11 3722.10 3717.92
I understand why this happens arithmetically. At roughly 2,200 observations per subgroup, BIC charges about 15 points for the two extra parameters and AIC about 4. The fit loss from fixing the seating standard deviations is around 7 to 12 points, which falls between the two penalties, so each criterion answers as its penalty dictates. What I am unsure about is what to do with that.
What else I have looked at. In Model A the standard deviation on seating level 2 is significant at the 1% level in all six subgroups, with point estimates between 0.64 and 0.84, so the heterogeneity being removed is not marginal. Model B also required non-zero starting values for the constants in one subgroup, where Model A converged from zeros throughout. Both considerations point toward keeping the parameters, but neither is decisive, and I am conscious of arguing my way toward the answer I happen to prefer.
My questions:
Which model should I do and based on what criteria?
Is there a settled convention on which criterion to prefer in choice modelling when the two disagree, or is it genuinely a matter for judgement in each case?
Where the disagreement is driven purely by the penalty on a small number of parameters, as here, does anyone use a formal test — a likelihood ratio test on the restriction, for instance — rather than relying on the criteria?
Does the significance of the standard deviations being removed carry weight in practice, or is that double-counting evidence the criteria have already used?