Friday, September 18, 2026

The Role of The Coach in Elite Equestrian Sport - Juniper Publishers

 

Physical Fitness, Medicine & Treatment in Sports - Juniper Publishers


Abstract

British Equestrian (BE) aims to develop a holistic coach education and certification program, moving away from traditional autocratic instruction in line with the United Kingdom (UK) Coaching Framework. This framework is based on generic coaching science research where the coach is cited as a pivotal aspect in developing sporting success. Theoretic knowledge suggests that the role of the sports coach is to develop the physical, tactical, technical and psychological attributes of the athlete and is responsible for the planning, organization and delivery of the training plan and competition schedule. However, there is no empirical evidence to suggest that is the role required in equestrian sport, as the rider often takes responsibility for many of these tasks. This research aimed to address the void in current knowledge by gaining an understanding of coaching in equestrian sport at the elite level, to improve coaching education systems through awareness of the role of the coach. A qualitative method using semi-structured interviews was used. A sample of elite coaches (N=3) and elite riders (N=3) were interviewed. Analysis of the transcripts revealed a total of 534 meaning units that were further grouped into sub-themes and general themes from the coaches’ perspective and the riders’ perspective. This led to the development of a final thematic structure revealing major dimensions that characterized coaching in elite equestrian sport. It was found that the riders at the elite level, coach themselves most of the time therefore can be considered as ‘self-coached’ athletes. However, they do use elite coaches in a mentoring and consultancy role, where they seek guidance from the coach on specific problems, to sound ideas off or to seek reassurance that what they are doing is correct. Findings from this research suggest that the rider-coach relationship at the elite level is a professional one, based on trust and respect, but not a close relationship, as seen in other sports. The results show the imperative need for the BE to educate coaches in coaching the self-coached rider at the elite level, particularly in terms of mentoring skills. As well as incorporating rider education aimed at developing the independent, self-coached riders.

Keywords: Coach; Elite; Equestrian sport

Introduction

Equestrian sporting origins are deeply rooted in military tradition, both in the development of the sport and the training of the horse and rider. Equestrian sports are unique as they test the rider’s mastery over the horse in terms of athleticism, control, and accuracy. British Equestrian (BE), the umbrella governing body of equestrian sport, represents 10 sports including the Olympic disciplines of evening, Show Jumping and Dressage. It works to promote the interests of 3 million riders and carriage drivers in the United Kingdom. The Federation is responsible for distributing government funds from UK Sport with an aim to win more medals on the world stage and to get more people participating in equestrian sports. BE places the coach as an integral part in achieving these aims. Substantial sports coaching literature has identified the importance of the role of the coach in sporting success as they are pivotal in the development of physical, tactical, technical, and psychological attributes of the athlete. However, equally important is the role the coach plays in the overall enjoyment, satisfaction and ultimately retention of people participating in the sporting activity. The basic role of any sport coach is to develop and improve the sporting performance of the team or individual. However, as participation in sport is usually voluntary, the experiences encountered can make or break a participant’s continuation in the sport. If the experience is not satisfactory, they are likely to leave the sport. The coach is therefore the key component to whether the activity is positive, and the quality of this experience depends on the coach’s value, principles, and beliefs.

It is acknowledged that the role of the coach is diverse and often not fully understood. Indeed, modern day coaching practitioners are not only responsible for directing practice and training sessions but also for the overall social and psychological well-being of their athlete both inside and outside of the sporting arena. Therefore, as [1] points out, to fully understand the role of the coach a critical analysis is needed of the nuances, actions, behaviour, and complexities used by sport specific coaching practitioners. This suggests that there needs to be recognition of the layers of skill and competencies required, how these interact with each other and how they impact on performance. Research has identified that successful coaches across a range of sports have several qualities in common: the ability to select the most important leadership behaviour; a personal desire to foster individuals’ growth; organizational skills in planning and preparation; a strong sense of goals, philosophy, and personality. This indicates that coaching cannot be viewed as purely a series of actions but a complex model of overlapping aspects. Research supports this by suggesting that coaches require additional skills above the technical knowledge of their sport and that these include pedagogical skills of a teacher, counselling skills of a psychologist, fitness training skills of a physiologist and the administrative and leadership skills of a business executive [2] also includes the role of a mentor and pillar of support to the athlete [3] clarifies the practice of coaching as a complex, dynamic, social domain and context dependent enterprise with often contradictory goals and values. Understanding these complexities is key to evaluating the purpose of the coach, needed in coach education to develop and improve coaching skills. The training of coaches is seen as essential to sustaining and improving the quality of coaching and on-going professionalism. Yet, [4] points out that currently coach development programs use a competency-based approach. This is, in fact, true of the BE who have developed a certification program in line with UK Sport’s United Kingdom Coaching Certification (UKCC). This is a move away from the traditional autocratic style ‘instructor’ to an ‘athlete-centered’, holistic approach to coaching. The traditional system, whilst acknowledged worldwide as a comprehensive program developing basic riding skills was an authoritarian approach to teaching riding. The syllabus was based on the ‘classical’ tradition of training horses based in the military past and lacked scientific validity. The current focus on the holistic approach and the development of an athlete, is supported in other areas of coaching, physical education and indeed education. A strong athlete-centered coaching approach emphasizing the development of self-confidence and belief in one’s ability is essential in making the correct decisions in a competitive situation. Producing ‘independent learners’ and ‘independent ‘decision markers’; may be key to the equestrian model as the rider is often considered as the ‘leader’ and non-verbal decision maker of the team.

Decisions made during riding must be made quickly and have dynamic consequences. The rider must calculate so many variables and translate them to effective communication with the horse. The rider needs to make the correct decisions at the right time and failure to make correct decisions can be catastrophic. This requires quick cognitive function, complex tasks, and choice reaction times. Indeed, it is this quick proprioceptive processing and effective decision making that are key to effective horse riding. This can only be achieved if the rider is empowered to be independent leaders that are confident decision makers. Olympic rider and coach, Phillip Dutton agrees: “You need to be strong and independent enough so you can ride without an instructor watching you all the time, holding your hand, and doing everything for you. Eventing is a sport where you are out there on your own especially cross-country phase. Your instructor may help you gain skills and improve your riding, but you must develop your mind and confidence so that when you are on course you can do it on your own [5].

To attain this holistic approach to coaching a substantial amount of literature has revealed the importance of the coach-athlete relationship and that the strength of this relationship should be based on closeness, co-orientation, and complementarity (3Cs). However, literature, mainly popular and non-academic, suggests that it is the rider that is responsible for the training of the horse in terms of fitness, skill, and technique as well as the management and wellbeing of the horse. They also assume the planning of competition schedules as well as most of the tactical and technical support, therefore it can be assumed that in part, they coach themselves. There is little empirical research to identify what the role of the equestrian coach is in this triad relationship. Therefore, the aim of the study was to gain an understanding of the role of the coach in elite equestrian sport to improve coaching education systems through awareness of the role of the coach. The first objective was to examine the relationship between coach and rider at elite level in equestrian sport providing empirical evidence to suggest that the rider is, in part, ‘self-coached’. The second objective was to identify what elite equestrian coaches believed their role is in the development of safe and effective riders.

Method

Sampling

The research question was developed with regards to elite equestrian coaches and riders as the study of their inferences could be drawn and applied to coach education. As there is no cohesive definition of an expert coach or valid ways to identify expertise, almost all research relies on years of experience or level of performance. As this study formulates primary research in the field of equestrian coaching and due to the large number of variables associated with the sport a top-down approach researching the elite level was chosen. Therefore, the principles of purposeful sampling were implemented using the following criteria: Three elite equestrian coaches were selected who were currently coaching on the BEF World Class Programme and working with GB senior team riders competing at the international level. A list of subjects that met the criteria, were drawn up and were approached based on location to researcher. One female coach and two male coaches were interviewed, one specialised in dressage, one show jumping and one in eventing, although all had had experience of coaching event riders. All were known in a professional manner to the researcher. Three elite riders were selected using the following criteria: member of the BEF World Class Programme and had represented Team GB at senior level during the past year. Three riders were selected from a possible 32 across the disciplines of eventing, dressage and show jumping. Two riders competed in eventing, one in the discipline of Show Jumping and were selected on location to researcher and were also known in a professional manner to the researcher.

Instrument

A semi-structured interview schedule that prompted responses to open ended questions about the roles and relationships of coach and rider was chosen as the method for this study. A semi-structured interview guide (Table 1 & 2) was developed to allow the interviewer to explore the relationships and roles within the coaching process. Participants were informed that there were no right or wrong answers, they were asked to take their time to respond to questions or to tell the interviewer if they did not understand the question. In addition to the semistructured questions specific probes were identified for each question and encouraged participants to elaborate on their responses.

Data collection

Coaches and riders were approached via phone call or email by the author and invited to participate in the study. After explanation of the aims and background of the study, interviews were arranged at a time and place elected by the interviewee. The interviews ranged from 25-55mins and were audiotaped with participants’ consent. The interviews were later transcribed verbatim into A4 single-spaced text.

Pilot study

A pilot study was carried out to assess the effectiveness of the questions selected. One elite coach was selected to be interviewed. Responses were forthcoming and met expectations. However, to validate the results further it was decided to also interview riders to gain their perspective of the role of the coach and the relationship they have with their coach. A further pilot interview was carried out with an elite rider to again assess the effectiveness of the questions selected.

Data analysis

Thematic Analysis was used as a systematic method for exploring the contents of the obtained data.

Ethics & Limitations

Ethical approval was granted by Hartpury University. The interviewees remained anonymous and confidential throughout the study, however, due to the high-profile nature of the sample it may be possible for people in the equine industry to identify subjects from their responses, however all efforts were made to retain anonymity and participants were given a letter and number as a means of identification through the study. Questioning topics did not cover intrusive or overtly personal subject areas. All participants were over the age of 18 and choose to take part under their own free will and were able to withdraw from the study at any time. Informed consent from the coach and riders were obtained prior to the interviews being recorded. The Data, both audio recordings and transcripts were stored in accordance with the Data Protection Act 1998.

The limitations of using a semi structured interview technique are that not all people are equally cooperative, articulate, and perceptive. The interviewer requires skill, not only to select and ask the appropriate questions clearly, but to gain the interviewees trust and confidence for them to elicit a full and honest response.

As a result, it is not a natural tool for gathering data as it requires interaction between two parties. However, due to knowledge level and experience of the researcher regarding equine performance these limitations are reduced.

Results

Analysis of the transcripts revealed a total of 534 meaning units that were further grouped into sub-themes and general themes from the coaches’ perspective (Table 3) and the riders’ perspective (Table 4). This led to the development of a final thematic structure revealing major dimensions that characterized coaching in elite equestrian sport (Table 5 & 6).

Discussion

Self-coach athletes

Analysis of data suggests that riders at the elite level in equestrian sport are in part ‘self-coached’. The results of the study show that riders attend a training session with a coach less than once a week and some as little as once a month. The remaining time the individual rider is responsible for all training decisions including horse selection; competition planning; implementation of periodization plans; management of support staff. Training decisions are made by the riders based on experience and knowledge of the individual horse. Therefore, in depth knowledge is needed by the rider in all areas. The development of the equestrian UKCC qualifications has incorporated the importance of planning for the coach but fails to acknowledge that it is the rider that plans the program. It was acknowledged by both the riders and coaches that consideration for the horse’s wellbeing and ensuring they were willing to do the task being asked was a key factor in their planning. This is because the equine has no concept of the goals involved. Coach C1 clarified this by stating “It is better to have the horse 90% prepared or 90% fit and 10% willing than have 100% prepared but you have no willingness”. This is in direct contrast to literature in other sports where being 100% prepared is necessary for sporting success. Therefore, any equestrian plan needs to suit the temperament of the horse. The results highlighted that knowledge and understanding of the psychology of the individual horse was significant and that as the rider has a close relationship with their horse, they are in the best position to make these training decisions. However, the equestrian UKCC syllabus does not incorporate any aspect of equine psychology or horse management, although these topics are still present in the other equine coach certification systems.

All riders in the study identified that how the horse ‘feels’ is the deciding factor in increasing intensity of training or increase in competition level, this suggests that there is an element of reflective analysis occurring and highlights the importance for developing and understanding this concept of ‘feel’. Yet there was no evidence to suggest riders use an objective analysis approach, this may be due to lack of knowledge by the rider and lack of research reaching the industry e.g., use of heart rate monitors, biomechanical analysis etc. There is also a lack of consistency both in literature and amongst the participants as to what actually constitutes ‘feel’. More research is needed to define this concept in equestrian sport and how it is developed to enable it to be taught or coached more effectively.

During their self-coaching training, all riders stated that they used outside observers either grooms, family members or friends to gain feedback. However, the quality of these observers is unknown. This is an area that requires focus and providing quality education to this support team needs consideration. A greater understanding of the self-coached riders is needed to fully feed into education programs for riders at all levels as well as an understanding of how the coach can support the self-coached rider optimally.

Coach-rider relationship

Current research in the field of coaching science recognizes that sport is not immune from the social world and that to examine the dynamic coaching process contextual social factors must be considered. Indeed, any activity that involves human beings is complex, interpersonal and that relationships are contested at levels of meaning, value, and practices. The relationship between the coach and the athlete is not an add-on or by-product of the coaching process but could be considered the foundation of coaching. Therefore, the coach-athlete relationship can be defined by mutual and casual interdependence between coaches and athlete feelings, thoughts, and behaviour, suggesting shared goals and values [6] proposed that successful coach athlete relationships are based on four concepts: closeness (trust and respect); commitment (shared goals and connection); complementarity (interaction that is co-operative and effective); co-orientation (acceptance of individual roles). However, previously there is no evidence to suggest these are used in equestrian sport.

The interpersonal relationship between athlete and coach plays a significant role in the sporting lives of the athlete and is likely to determine satisfaction in the sport, self-esteem and confidence and ultimately successful sporting performance.

However, knowledge and understanding of these relationships remains limited at both the theoretical and empirical level. This study elucidated that the relationship between the coach and rider was a professional one that could be described as reciprocal, yet asymmetrical characterised by a unique relationship between individuals and depended upon experience and age difference. This study revealed riders sought a coach that was approachable, that they felt comfortable discussing ideas with and that they trusted and respected. This view is supported in other sports; however, the riders did not describe the relationship as ‘close’. Yet closeness is considered by Jowett’s 3Cs as a key component of the in successful coach-athlete relationships. This may be in part due to the limited contact riders have with their coaches. Contact was mixed amongst the subjects; one respondent only saw their coach during periodic World Class training which may explain why their relation was not considered ‘close’. Further research is needed to fully quantify the ‘norm’ for contact time with coaches across equestrian sport at varied rider levels.

Respect and trust

The emerging themes that were expressed by the rider participants show some commonality in their lower order themes, that they desire a coach that they trust and respect. All participants claimed that this respect was generated from the coaches’ own riding experience and level in which they had competed. It was felt that this was needed to not only have credibility but also to have the knowledge of riding a variety of horses at the elite level and to have the appropriate repertoire of training solutions. The ability to have the concept of ‘feel’ of the individual horse was also deemed important. One participant commented that: “It is also really useful for them to get on the horse so they can feel what I feel” R1. This suggests that equestrian coaching is largely experienced-linked and situation-specific base, like that required in the sport of sailing. Such an important statement is worthy of further investigation as other equestrian coaching qualifications include riding tests as part of their qualification curriculum, whereas the UKCC does not. Interestingly the elite coaches acknowledge the advantage of riding experience for a successful coach but felt that this was not necessary. This is supported in other sports where the best coaches are not always elite athletes but have had experience of competing just below the elite level, this may well not be applicable in equestrian coaching at elite level.

Mentor relationship

Riders in this study referred to their chosen coach when they needed advice or mentoring. One way by which the riders identified this mentoring relationship was that they used a coach as a sounding board for ideas. This was, in part, used to gain confidence and reassurance in the knowledge that what they were doing was correct. They also used the coach as a mentor when they had a particular problem or needed a fresh approach to a particular horse or situation. Whilst BE acknowledges the importance of mentoring skills within coaching, it fails to clarify what these skills actually are, yet it can be accepted that it is a form of supported learning through social interaction. Evidence from these interviews suggests this is achieved through a shift between support and challenge.

Facilitator in the development of safe and effective riders

[7] when analysing relationships within the caring professions, identified good mentors as challenge givers, the collective viewpoint expressed by the participants indeed concurs with this within this study. Emphasis was placed on the element of challenging the rider. This may be since in the remaining time the riders are self-coaching and may not be motivated or confident to push themselves outside their comfort zone. C1 expressed this view “that they go over what they are comfortable doing and that they are good at” C1 29-30. Within this study the findings concluded that the elite equestrian coaches facilitated this challenging environment by setting up exercises that allow riders to experiment, for example, different approaches to jumping a combination. This suggests a move away from skill practice to development of perception and decision-making processes. This allowed the riders to make mistakes and learn from these mistakes. The coaches in this study stated that they achieve a learning experience by discuss those mistakes, getting them to think how they would ride the exercise differently and creating awareness of feel in relation to position. This cognitive action through a guided discovery approach achieves an empowerment process.

Similarly, to other sports, feel or body awareness in equestrian sports is achieved through drills and repetition. Riders felt this fed into positively developing their own confidence, improving their cognitive awareness and automatic decision-making processes. More research is needed in this area to understand which exercise or drills are the most effective in developing this aspect within equestrian coaching.

Recommendations

The results from this study provide substantial evidence for the need to incorporate the topic of coaching the self-coached rider into equestrian coaching education systems at elite level. Coaches should be aware of the demands and limitations of coaching the self-coached rider and appreciate the importance of their role in the successful outcome of the horse/rider dyad in equestrian competition. BE should highlight and develop the role of the coach as a mentor to self-coached riders at the elite level. Amalgamation of both phases of this study combined with the themes that emerged from the interviews provides the following recommendations:

a) Role of the coach in Equestrian UKCC education should be clearly identified

b) Further development of mentoring skills of coaches

c) Identification of techniques that facilitate the development of ‘self-aware’ and ‘self-reliant’ effective decisionmaking riders

d) Development of rider skills to self-coach in terms of planning and implementing training, developing all areas of psychology, equine psychology, injury prevention etc. and analysis

e) Increasing the use of tools to enable the self-coached rider to analyse performance in both training and competition environment

Limitations of Study and Future Research

It is important to highlight the limitations inherent in the study which must be considered against the results that emerged. The sample size used in the study was small (coaches N=3, riders N=3) however, the selection criteria was carefully applied and even though the sample size was small it could justifiably be seen as offering expert opinions therefore, the findings are directly applicable to elite coaches and riders. Future research is required with differing levels of riders and equestrian coaches working to provide validity across all equestrian participants. More indepth research is indicated investigating individual equestrian sports in greater detail to examine any differences that may arise in each discipline. As expected, with any attempt to summarize or condense findings from the semi-structured interviews, not all participants were as forthcoming as each other and did not respond in the same way or to the same extent to the identified themes. The practical coaching processes were not quantified, therefore this study relied on the participants perception of coaching and the role of the coach, whilst this is a legitimate form of qualitative research the study could have included coaching observations. using video analysis to evaluate the coaching process and identify evidence of the coaching roles displayed in practice [8-54].

To Know more about  Journal of Physical Fitness, Medicine & Treatment in Sports

Click here:   https://juniperpublishers.com/jpfmts/index.php

To Know more about our Juniper Publishers

Click here: https://juniperpublishers.com/index.php



Tuesday, September 15, 2026

Frequency of Visiting a Doctor: A right Truncated Count Regression Model with Excess Zeros- Juniper Publishers

 

Biostatistics and Biometrics - Juniper Publishers


Abstract

Count response variables are frequently encountered in medical data, which calls for the use of count regression models. In this study, we introduce the hurdle Conway-Maxwell Poisson (HCMP) regression model where the outcome variable is the number of doctor visits, complicated by excess zeros and over-dispersion from troublesome extreme values. A truncation approach is proposed to handle extreme values, leading to the definition of a truncated HCMP (THCMP) model. Parameter estimates are derived using maximum likelihood. Results of a case study on a RWM dataset investigated effects of response truncation at 6.65, 3.08 and 1.75% for the THCMP and truncated hurdle Poisson (THP) models. In a simulation study, responses were generated from a mixture of HCMP (50%) and HP (50%) probability models. THCMP and THP model performance was compared with respect to parameter estimation bias, goodness-of-fit and outcome estimates for truncation levels of 5 and 10%. As measured by AIC, the THCMP model exhibited better goodness-of-fit at all truncation levels compared to the THP model. Estimation bias increased with higher truncation levels for both models, but to a lesser degree for the THCMP model.

Keywords:Hurdle model; Conway-maxwell poisson; Over-dispersion, Parameter estimation; Model selection

Abbreviations: HCMP: Hurdle Conway-Maxwell Poisson; THP: Truncated Hurdle Poisson; CMP: Conway-Maxwell Poisson; GSOEP: German Socioeconomic Panel; LL: Log-Likelihood; AIC: Akaike’s Information Criterion; BIC: Bayesian Information Criterion; TP: Truncated Poisson; TCMP: Right-Truncated CMP; HNB: Hurdle Negative Binomial; HGP: Hurdle Generalized Poisson

Introduction

Health care is one of the most important factors in human life. Good health care is a major contributor to quality of life, and ready access to a physician is an important component of a good health care system. The number of doctor visits for a household over a fixed interval is a useful metric for studying factors that affect physician accessibility. In this study, we introduce a new regression model for studying count data occurring in medical studies. The count variable in this study is number of doctor visits over a fixed time period.

There are numerous publications describing applications of count models to healthcare demand data, and applications of negative binomial models in particular. The negative binomial model has been applied to cross-sectional data, and in econometric models to analyze cross-sectional data with multiple outcomes per observation [ 1]. Modelling count data using a random effects negative binomial regression model is discussed in [ 2], and application of the negative binomial hurdle model to physician visit data is demonstrated in [ 3].

The Conway-Maxwell Poisson (CMP) distribution-a generalization of the Poisson-was introduced in [ 4] with applications to queues and service rates. The CMP distribution belongs to the exponential family and the two-parameter power series family of distributions. In the 50 plus years since its introduction, the CMP model has not been widely employed; however, a revival has arisen of late from a recognition of its utility in fitting discrete data [ 5]. The CMP distribution has two parameters and can handle both under- and over-dispersed data. This is in contrast to the commonly used negative binomial model which can only handle overdispersion.

We found several studies describing the properties of the CMP model with a variety of applications. The CMP distribution was applied to model timing of bid placement and the extent of multiple bidding in online auctions [ 6]. A Bayesian analysis of the CMP distribution is discussed in [ 7] and conjugate priors for the distribution parameters are derived. A flexible cure rate survival model is expanded to follow the CMP distribution in [ 8]. The joint generalized quasi-likelihood estimating equations are compared to the marginal equations in a CMP generalized linear model describing the number of car breakdowns in [ 9]. The structural properties of the CMP distribution, including moments and probability generating function are derived in [ 10]. Notwithstanding the many published applications, we have yet to find the CMP distribution used in a medical context-which is the motivation for this paper.

In many real world applications, the problem of excess zeros is encountered-the actual zero frequency outcome is higher than that predicted by the theoretical model. When this is the case, a hurdle model may be used that models the zero outcome separately from the non-zero outcomes [ 11]. In a hurdle model two densities are used, one that generates the zeroes, and another called the zero-truncated density that generates the positive values. The finite mixture distribution generated by combining two densities is discussed in [ 12]. The mechanisms by which excess zero frequencies occur for various types of count data, and how the zero-inflated Poisson model has application to such data in a medical context are described in [ 13]. Excess zeros in count data with application to public health employing a likelihood ratio test is addressed in [ 14].

Another aspect of our paper focuses on extreme values in the context of the CMP model and the adverse effect of troublesome ‘outliers’ on estimates of the CMP distribution mean and variance. The effect of outliers is to inflate the variance to make it larger than the mean, which is the definition of over-dispersion in the CMP model. One approach for reducing over-dispersion is right truncation, where values greater than a fixed constant are removed from the sample. A right-truncated Poisson regression model for handling over-dispersion is discussed in [ 15]. Applications of hurdle models with right censoring are the hurdle generalized Poisson regression model and the hurdle negative binomial regression model applied to the number of fish caught by fishermen at a state park, where the response was right censored [ 16,17].

The main focus of this study is on regression analysis based on CMP distribution. CMP has two parameters and this feature makes the distribution more flexible compared to Poisson model. Negative binomial (NB) model is a competitive model for CMP, however the dispersion parameter in NB model can only deal with over-dispersed data. CMP model is more flexible in that sense and can handle both over- and under-dispersion scenarios. The application part of this study (including read data example and simulation study) illustrates the performance of CMP model over alternative models under both under- and over-dispersed data.

The novel contribution of our paper is the introduction of a hurdle model based on the CMP distribution, the truncated hurdle Conway-Maxwell Poisson (THCMP) distribution, which can handle excess zeros in right-truncated count data. We illustrate the THCMP model in an application involving an analysis of the number of doctor visits over a fixed interval and compare it with the truncated hurdle Poisson (THP) model—a less suitable, but possible choice among presently available alternatives. In section 2, we describe the health care data set and variables for a case study analysis. Inasmuch as we have found no published study applying the CMP distribution to an outcome in medicine or public health, our paper is unique in this regard. In section 3, the THCMP regression model for right truncated data is introduced and parameter estimates derived. A case study analysis of the THCMP model is presented in section 4 along with a description of the methodology for a simulation study in which the THCMP model is compared to the truncated hurdle Poisson (THP) model on both over- and under-dispersed data. In section 5, we discuss results of a simulation study and evaluate the performance of the THCMP regression model versus the THP model.

Description of RWM health care data

The RWM data set [ 18] used in this study is taken from the German Socioeconomic Panel (GSOEP). The GSOEP, conducted by the German Institute for Economic Research in Berlin, surveys a representative sample of East and West German households. Researchers have recently used this cross-sectional data set to evaluate performance of their proposed count regression models [ 18,19]. The RWM data set is an unbalanced panel survey of health care utilization of 27,326 German individuals. The frequency table for the number of visits to the doctor is given in Table 2.



The dependent variable in our analyses, DocVis, is a count variable—the number of visits to a doctor (including dentists) during a fixed time interval. The RWM data set is viewed as a cross-sectional dataset in our analysis [ 18,19]-the outcome is not time varying, as contrasted with [ 18]-and we assume that counts among study subjects are independent. The particular irregularities/anomalies of this dataset make it a good candidate for illustrating the unique features of the THCMP model. In Table 1, the mean and variance of DocVis are 3.18 and 32.37, respectively, which indicates substantial overdispersion in the data; minimum and maximum values are 0 and 121, respectively. In addition, the frequency of the zero response in DocVis is higher than expected (median=1, mode=0) (Table 2).

The explanatory variables consist of socioeconomic characteristics and demographic variables. All count regression models on DocVis were fitted as functions of sex (1=female, 0=male), age (years), education (years of schooling), marital status (1=married, 0=single) and children in the household (1=children present). Independent variables are summarized in Table 1 which shows that 48% of visits are by females, average age is 43.5, children are present in 40% of households, average years of schooling is 11.3, and 24.1% of respondents are single. The base case count model used in the analysis included the following variables in addition to the constant term:

The frequency distribution of DocVis is shown in Table 2. According to the percentage of zeros in the response variable (37.1%), there is an excess of zeros. In addition, the 95th percentile is 12 which means that there are some extreme values in the sample. It is apparent from the histogram in Figure 1 that the zero count frequency of DocVis exceeds that expected in a Poisson distribution.


Methodology

In this section, the right-truncated hurdle Conway-Maxwell Poisson (THCMP) regression model is introduced for handling count data with excess zeros and right-tail data truncation. Parameter estimation and the goodness-of-fit statistics are discussed.

1.1. The model

Let the response variable *,1,,iYin= be the number of visits to a doctor over a fixed time period. The HCMP regression model ()*,,,iifvyλ is given by

Where ()*iiEYλ= of a Poisson distribution associated with observation , and 0v≥ is the dispersion parameter. The CMP regression model can handle both over-dispersion ()1v< and under-dispersion ()1,v> and when 1,v=the probability function (1) reduces to a Poisson model. A geometric distribution is obtained from (1) when0v= and 1.iλ< When v→∞ in (1) with probability ,1iiλλ+ the result is a Bernoulli distribution.

In many practical applications, it is common to assume that the parameter iλ depends on a vector of explanatory variables

variables are commonly incorporated in the context of a log-linear model, where log indicates the base e or natural logarithm,

The'jsβ are coefficients of the explanatory variables in the regression model and m is the number of explanatory variables.

vector of unknown parameters. In this model set up, the non-negative function 0w is modeled using a logit link function.

The moments of the HCMP distribution are obtained as

Now, we can define the right truncated hurdle Conway-Maxwell Poisson regression model as

Where t is the truncation point for .iy This means that we truncate the response variable when ,iyt> leading to the definition of B as

Thus, the log-likelihood function for the HCMP model with right truncation can be written as

Where k is the number of observations after truncation.

Parameter estimation

In this section we obtain parameters estimates using maximum likelihood. The likelihood equations for estimating ,rtβδ and v are obtained by taking the partial derivatives of (4) and setting them equal to zero yielding

These partial derivative equations cannot be further simplified. Calculating the Hessian matrix directly is computationally laborious, and so the Conjugate Gradient Optimization method implemented in SAS was used to numerically obtain the Hessian variance-covariance matrix. In approximating standard errors of the parameter estimates, the Hessian matrix must be computed at least once, regardless of the optimization technique.

The Fisher information matrix for the THCMP regression model is obtained as

The elements of the Fisher information matrix are available in the Appendix.

Model selection and test for dispersion

Goodness-of-fit statistics for the THCMP model are based on the deviance statistic, defined as

the model likelihood function evaluated at μ and ,y respectively. The log-likelihood function is defined in equation (4).

The deviance statistic can be approximated by a chi-square distribution when 'sμ is large. In the application section, we use 2,LLAIC− and BIC to compare the different regression models in terms of goodness-of-fit. For all of these statistics, a smaller value indicates a better fit.

From section 3.1, it is obvious that the THCMP model reduces to THP model when 1.v= To assess the adequacy of the THCMP model over the truncated hurdle Poisson model, we test the hypothesis

The purpose of (5) is to evaluate the significance of the dispersion parameter. It follows that the THCMP model should be used instead of the THP model whenever 0H is rejected. To test the null hypothesis 0H in (5), the likelihood ratio statistic can be used. An alternative statistic for the parameter v is the asymptotic Wald statistic where the dispersion parameter is calculated after fitting the THCMP regression model.

Results

Case study

In this case study using the RWM data set (n=27,326), the THCMP regression model is used to model the number of doctor visits (DocVis) per patient over a period of three months as a function of the independent variables sex, age, children, education and married. The THP model is also considered as an alternative model in the analysis of the RWM. The THCMP model will be compared to the THP model relative to parameter estimates, standard errors, goodness-of-fit statistics and accuracy in modelling the response.

Five independent variables are used in the model and all are incorporated into both the logit and non-logit parts of the model. Therefore, the link functions can be written as

plus the dispersion parameter (when used) will be estimated using the ML method.

Three truncation points, t=10, t=15, and t=20 are employed in comparing the effects of truncation, and correspond to truncation percentages of 6.65, 3.08 and 1.75, respectively.

Parameter estimates for the THP and the THCMP model were obtained for the specified truncation points and are summarized in Table 3. Link functions are obtainable from the estimates shown in Table 3. For example, the respective log and logit link functions for the THCMP regression model for truncation point 110t= is

Using the log link function coefficients in Table 3, one can see the estimated change in DocVis per unit change in each independent variable, all others being held constant. To illustrate, the positive 1β coefficients corresponding the variable sex in the HP model in Table 3 for the three truncation times are 0.0997, 0.11 and 0.0935, which indicate higher values of ()logDocVis for females compared to males for all truncation points. For the THCMP model, decreases in ()logDocVis of 0.0158 and 0.003 for females versus males are predicted for 215t= and 320,t= respectively. For a one-unit increase in age, expected ()logDocVis is estimated to increase by ~0.01 on average using the THP model and ~0.003 using THCMP model for all truncation points. Thus, older patients are predicted to have a higher number of doctor visits per unit time. ()logDocVis is estimated to be lower for households with children relative to those with no children for both THP and THCMP models at all truncation points—with the exception of THCMP when 215.t= For a one-unit increase in the number of years of schooling, the estimated change in the number of doctor visits decreased for all truncation points. ()logDocVis showed more frequent doctor visits for married versus single individuals for all truncation points.


The THCMP regression model indicated overdispersion in the RWM data. Dispersion parameter estimates corresponding to truncation percentages of 6.65,3.08 and 1.75 were 0.108,v= 0.069v= and 0.0621v= respectively, where 1v< indicates overdispersion.

THP and THCMP regression model coefficients for the RWM analysis based on the logit link function are also given in Table 3. These coefficients correspond to the excess zeros component of the models. Both models show a negative effect of sex for all truncation percentages, indicating a higher rate of zero doctor visits in males than in females. The log odds of excess zeros increases for each unit decrease in age in both the THP and THCMP models at all truncation points. This means that zero doctor visits were increasingly more likely with aging. A positive coefficient for children in both THP and THCMP models indicated that households with children exhibited a higher rate of zero visits to a doctor compared to households without children for all truncation percentages. The log odds of excess zeros increased for each unit increase in the number of years of schooling for all truncation points for both regression models. This indicates that fewer years of schooling were associated with higher odds of zero visits to a doctor. Both THP and THCMP models resulted in negative coefficients for the married variable for all truncation points with the exception of the 320t= truncation point for the THCMP model. So, generally, single individuals exhibited higher rates of zero doctor visits than married individuals.

Table 4 compares the THP and THCMP regression models on three goodness-of-fit measures: log-likelihood (LL), Akaike’s Information Criterion (AIC) and the Bayesian Information Criterion (BIC). The right-truncated Poisson (TP) and right-truncated CMP (TCMP) regression models are included in Table 4 as special cases to investigate whether a zero frequency of 37.1% should be considered an excess zero scenario. We were also able to investigate whether the TP and TCMP regression models (without excess zero scenario) were able to fit the data as well as the hurdle regression models. The THP and THCMP models demonstrated better goodness-of-fit than the TP and TCMP models on all measures (smaller is better). The THCMP regression model exhibited superior goodness-of-fit compared to the THP regression model for all truncation levels in the RWM data analysis.


Table 5 shows that the THCMP regression model exhibited a better fit to the RWM data versus the THP model based on the predicted versus observed frequency count criterion for all truncation points. The RWM dataset observed zero frequency was 10135. Predicted zero frequencies for THCMP/THP at the truncation points were 1:10118/10012;t 2:10319/9971;t and 3:9626/9008.




A subset of RWM data where under-dispersion is present was considered. We have focused on the individuals with low health satisfaction score (<4 out of 10) and split up the data set by marriage status to come up with two under-dispersed scenarios. The married covariate was excluded from the list of independent variables and other covariates were kept in. TP and THCMP models with a right truncation point of 4 applied to the data. The goodness-of-fit statistics (-2LL and AIC) for THCMP/THP were 653.2/686 and 675.2/706 for unmarried individuals, and 1847.5/1908.9 and 1869.5/1928.9 for married individuals. The dispersion parameter of THCMP model was significant in both scenarios (4.12 (3.02, 5.33) and 3.42 (2.77, 4.07), p-value<0.001).

A Simulation study

We conducted a simulation study to assess and compare THP and THCMP regression model performance. Goodness-of-fit was measured using the Akaike information criterion (AIC). Data were simulated with 50% of responses generated by the HCMP regression model and 50% by the HP model. Simulation models

erated from a uniform distribution on []0,1. A sample size of 200n= was used in conjunction with varying proportions of zero outcomes and 0w set to values of 0.2, 0.3 and 0.4. Truncation percentages were set at 5% and 10% of the simulated distribution tail, although actual simulated truncation percentages were not always exactly equal to 5% and 10%.

The six working simulation models, differing in their coefficient parameters, are shown in Table 6. Three are mixtures of under-dispersed (υ>1) HCMP regression models with the HP regression model. These models generate count data outcomes of 0, 1, 2, and 3 for the most part, resulting in short-tailed count frequency distributions. Conversely, the remaining three models are mixtures of over-dispersed (υ<1) HCMP regression models with the HP regression model, which generate more extended right tails in the response distribution. Using the three combinations of coefficient parameters, the model with 01,β= 11.5,β= 20.9,β= 01a= generates long right-tailed count distributions with high count outcomes; the model with 01,β= 11,β= 21,β= 02a= generates distributions with tails of intermediate length; and the model with 01,β= 10.8,β= 20.8,β= 00.3a= distributions with relatively short tails.

The number of replications was set at 1,000, which was sufficient for our purposes. Additional replications would have unnecessarily increased the computational burden. The simulation was programmed in FORTRAN. Maximum likelihood estimates were obtained via numerical maximization using the simulated annealing algorithm [ 20].

Simulation results are summarized in Table 7(i) and (ii) where the values given are the average values of 1,000 replications. From Table 7(i), the average AIC for the THCMP regression model was significantly lower than that for the THP regression model for simulation models (a) to (c). When the truncation percentage was increased to 10% as in Table 7(ii), the average AIC difference was smaller as compared to Table 7(i). Nevertheless, the THCMP regression model performed slightly better for model (a) and definitely outperformed models (b) and (c). For model (d), the average AIC difference between the THP and THCMP regression models was less than 2, regardless of the proportion of zeros and the percentage of truncation. For models (e) and (f), the THCMP regression model was slightly better than the THP model with 5% truncation in the tail, where the AIC average difference was greater than 2.

Discussion

In this paper, we introduce the THCMP regression model and illustrate its application in an analysis of the RWM data in which the outcome variable of interest is the number of doctor visits occurring in a fixed interval. We show how the THCMP model can be used to handle dual data anomalies of excess zeros and extreme values in a count response variable. Parameter estimates and standard errors for the TCMP model—and the alternative THP model-were obtained for selected data truncation levels using ML estimation. Covariate effects were interpreted in the context of model link functions.

A comparison of goodness-of-fit of the THCMP and THP models to the RWM data, as assessed by -2LL, AIC and BIC, showed better performance for the THCMP model at the three truncation levels studied (6.65%, 3.08% and 1.75%). The percentage of zeros in the response variable of the RWM case study was 37.1%—which is cited in the literature as a threshold for excess zeros [ 18, 19]. The results showed that a right truncated hurdle model can fit these data with inflation at zero better than a simple right-truncated model where excess zero part of the model is not taken into account. We also examined goodness-of-fit for the right truncated Poisson and CMP models as well as the excess zero models in the RWM case study analysis.

In the RWM analysis, the THCMP model exhibited better agreement between observed and predicted zero frequency counts than the THP model for all three truncation points. However, greater truncation levels resulted in greater bias in estimating the zero frequency for both models, although to a lesser degree for the THCMP model. In addition, the THCMP model generally exhibited better goodness-of-fit in modelling counts greater than zero. Exceptions favoring the THP model were a few cases involving frequency estimates for 2 and 7 at 6.65% truncation, and 8 and 9 at 3.08% and 1.75% truncation. The case study analysis results suggest advantages of the THCMP regression model compared to the THP model for analyzing ‘real life’ count data when there are both an excess of zeros and extreme values in the observed response.

In the under-dispersed subset of RWM data sets, we have tried hurdle negative binomial (HNB) and hurdle generalized Poisson (HGP) models with right truncation approach to investigate the performance of HCMP over HNB and HGP. HNB model was not converged because the final Hessian matrix, though full rank, had at least one negative eigenvalue, and therefore the second-order optimality condition violated. This is expected as NB model can handle over-dispersion scenario not under-dispersion. HGP model also was not converged as the final Hessian matrix was not positive definite and therefore the estimated covariance matrix was not full rank and may not be reliable. Hence, the HCMP model outperform HNB and HGP when the under-dispersed right truncated outcome has excess zeros.

In the simulation study, mean AIC for the THCMP model was significantly lower than mean AIC for the THP model (AIC mean difference >11) in under-dispersed scenarios with 5% tail truncation, indicating substantial improvement in fitting the data [ 21]. At the 5% truncation level, a clear lack-of-fit of THP is exhibited for simulation models (a) to (c), indicating under-dispersed scenarios, compared to THCMP. However, at 10% truncation, the mean AIC difference between models was smaller in under-dispersed scenarios. In the simulation study, the THCMP regression model performed slightly better than the THP model in the short-tailed data scenario (model (a)) and clearly outperformed THP in the medium- and long-tailed data scenarios (models (b) and (c)). In an over-dispersed scenario with short tail (model (d)), the average AIC difference for the THP and THCMP models was less than 2 regardless of the proportion of zeros or the truncation percentage, implying similar performance in this case. For over-dispersed scenarios involving medium and long tails (model (e) and (f)), and less than 5% truncation, the THCMP regression model performed slightly better than the THP model (AIC mean difference >2). In short, the comparison of average AIC for the THCMP and THP regression models indicated superiority of the THCMP to the THP model for outcome data with lower than expected proportions of zeros, lower percentages of tail truncation, and consists of mostly low values of the response variable.

A strong point of the simulation study is that the CMP model, unlike more frequently used models such as the negative binomial model, is more flexible and able to handle under-dispersion as well as over-dispersion. The negative binomial model was not entertained as an alternative model in this study because the NB model cannot accommodate under-dispersed data. The THCMP model clearly outperformed THP model for under-dispersed scenarios. Therefore, the THCMP regression model would be expected to provide better outcomes in terms of the parameter estimates and goodness-of-fit statistics when data are under-dispersed—even when compared to some alternative models such as hurdle negative binomial model.

In summary, we introduced the truncated hurdle Conway-Maxwell Poisson regression model. We carried out a simulation study, based on data generated from a mixture of HCMP (50%) and HP (50%) probability models, which showed the THCMP regression model accommodated various degrees of truncation and anomalous zero frequencies better than the THP regression model. We recommend the THCMP regression model as a flexible distribution for analyzing count data exhibiting the dual anomalies involving zero frequencies and a low level of tail truncation in handling over- and under-dispersion.

To Know more about   Biostatistics and Biometrics Open Access Journal

Click here:   https://juniperpublishers.com/bboaj/index.php

To Know more about our Juniper Publishers

Click here: https://juniperpublishers.com/index.php




The Role of The Coach in Elite Equestrian Sport - Juniper Publishers

  Physical Fitness, Medicine & Treatment in Sports - Juniper Publishers Abstract British Equestrian (BE) aims to develop a holistic coac...