Assignment I for Biostats Course VHM 801 at AVC - Fall Semester 2026

The assignment is worth 10% of the final course mark. Please be aware that by handing in the home assignment you implicitly acknowledge to have read and accepted the instructions for home assignments as described on the VHM 801 homepage.

This assignment is based on data collected by researchers from the Department of Companion Animals at the AVC. Their study involved blood samples from 30 healthy dogs of the same breed, taken at baseline as well as immediately prior to and post an intervention (to the dogs). A range of different measurements taken at the baseline visit were used to screen a larger pool of dogs for any condition that would exclude them from participating in the study. The intervention took place less than a month after the baseline visit. Our interest here is in analysis of the blood samples obtained for a specific biomarker. Three different analysis tool kits, here denoted as methods A, B and C, were applied to suitable subsamples. For reasons of no interest for this assignment, the blood samples prior to and post the intervention were not analyzed by method C. Therefore, a total of seven measurements are available for each dog, in the data files (in Minitab and comma-separated formats) labelled as:

in a notation where the prefix (A,B,C) indicates the method, and the suffix (base,pre,post) indicates the time of blood sampling. In addition, the data includes a variable (dog_no) labelling the dogs from 1 to 30.
Note: The values of the measurements in this dataset have been altered to protect the ownership of the data; these modifications have not substantially changed the relationships between the variables.

The home assignment has five questions (a)-(e) which should all be answered. An additional question (f) is optional, and a correct answer to this question can offset any loss of points on questions (a)-(e).

  1. Characterize the study type (e.g., experimental or another type), and determine the type(s) of variables present in the dataset (e.g., categorical nominal or ordinal, or quantitative discrete or continuous).

  2. Assuming that the study targeted 30 dogs, discuss how you would select the 30 dogs from a larger pool of dogs (say, 50 dogs) that passed the initial screening for the exclusion criteria. If you propose a random selection, describe how it will be carried out, in enough detail for someone else to reproduce your method. If you propose a non-random selection, discuss whether your selection could affect the population of dogs your study is intended to be representative for.

  3. Select one of biomarker variables (as you want), and carry out a descriptive analysis of your selected variable, including both a graphical representation and descriptive statistics. Choose the graphical representation and the statistics you find most useful to show its distribution. Comment specifically on the distribution's center, spread and shape. If your descriptive analysis identifies any "suspected outliers", discuss whether these should be considered as truly outlying observations, in the sense that they don't really belong to the distribution, or whether they should be considered as part of the distribution. (Hint: For such a discussion, it can be useful to look at all values obtained for the dog(s) in question.)

    Discuss also whether the distribution in your view is approximated well by a normal distribution. Describe carefully how you quantitatively assess the agreement of a variable with a normal distribution. If your variable is not approximated well by a normal distribution, describe how the distribution seems to differ from a normal distribution and explore briefly whether an improved approximation by a normal distribution can be obtained for either square-root or logarithmic transformed values. (Hint: You can use Minitab's built-in Calculator to compute the values on transformed scales.)

  4. Compute the change in the biomarker from before to after the intervention (for an analysis method of your own choice, not necessarily the same as you used above), and carry out a descriptive analysis similar to part (c) for this new variable. For this part, however, an exploration of transformed scales is not required.

  5. A reference interval for biomarker values of healthy dogs obtained by method A has been given as 17-140 (adjusted to the modifications applied to the data). Assume this reference interval was constructed to include the middle 95% of the distribution, as is commonly the case. The aim of this question is to compare the observed values for method A with the reference interval. Include only the values from one of the three samples per dog in the assessment; state and justify the sample you choose.

    1. If the reference interval matches the population represented by the study dogs perfectly, what will be the (theoretical) distribution of the number of dogs whose values (in your chosen sample) fall within the reference interval? Calculate the mean and standard deviation of this distribution.

    2. Determine the proportion of observed values (in your chosen sample) that lie within the reference interval and compare with the expected value. Compare also the number of observed values within the reference interval with the distribution from (i): is the observed number of dogs within the reference interval extreme in this distribution? (Hint: A graphical display of the distribution may be helpful.) Draw conclusions about how well the observed values compare with the reference interval.

    3. Additionally, compute also the probability for the reference interval when assuming the biomarker values to follow a normal distribution with its mean and standard deviation calculated from the data. You may perform the calculation on original or transformed scales, of your own choice. Compare your result with the proportion computed under (ii) and describe any differences you see.

  6. (Optional question) Different scientific questions can be explored with these data. For example, one may want to assess the effect of the intervention, describe how well different samples from the same dog analyzed by the same method agree, or compare the values obtained on the same sample by different methods. Select one such question (not necessarily among the examples provided) and determine a graphical display that in your view provides relevant information about the question being considered. Describe how you constructed the display and interpret carefully what the graphical display tells you about the question. Note that you are not required to perform any statistical analysis (because such analyses will most likely be beyond the scope of what has been covered in the course so far).


Henrik Stryhn (hstryhn@upei.ca) 2026-09-30