Estratto del documento

Name formula meaning

Name formula meaning 1 ∑ = ̅ = The Mean =1 Mean of a ∑ ⋅ = Values (ex: employee) Frequency =1 = ̅ = = Absolute Frequencies (ex: n° of shops) Distribution =1∑ Weighted ̅ = = ex. CFU =1∑ mean It is the value the stay in the central position in the sorted list of all the values.

The median

The Median x  x  x  ...  x  x      min 1 2 n max

The mode

The Mode It is the most frequent item.

The range

range  x  x The Range max min It is the width of the interval that contain Interquartile 50% the values of the distribution. (central dQ  Q  Q3 1 ones). – removing the highest and the Range lowest 25%.

The variance

The Variance (Deviation it computes the distance of each variability 12 2 2( )∑ = − ̅ ≥ 0=1 from the mean from the mean)  The The sum of squared deviation, the 2() )∑(= − ̅ numerator of the variance Deviance =11.

The standard deviation

The Standard Deviation 2) squared root of the variance ∑( = √ − ̅ Deviation =1 The ratio between the standard dev. and The Coefficient the mean, multiplied 100 = 100 ̅ > 0 (when you want to compare the variability of Variation of different sample).

The variance from a frequency distribution

The Variance 1 from a 22 ∑( ) = − ̅ ⋅ = Absolute frequencies Frequency =1 Distribution ̅ If a quantitative variable X as mean and standard deviation σ, it is possible to ( )− ̅ Standardised obtain its standardised values. The = values distribution of Y has zero mean and standard deviation equal to 1.

Covariance

Association between two quantitative variables:

  • The Cov > 0 Positive Cov )( ̅)∑((, ) = = − ̅ −
  • Cov < 0 Negative Cov
  • Cov = 0 Null Cov

Covariance =1 Only for LINEAR relationship.

Linear correlation “Pearson correlation coefficient”

(, ) = ( ) = = Pearson correlation coefficient, only for quantitative variables:

  • = −1 perfect negative linear =1∑ ( )( ̅)− ̅ − relation Correlation =
  • −1 < < 0 negative linear relation “Pearson ̅=1 2 2( ) ( )√∑ ∑− ̅ ∙ −
  • = 0 absence of linear relationship =1 correlation
  • 0 < < 1 positive linear relation coefficient”
  • = 1 perfect positive linear relation −1 ≤ ≤ 1.

Confidence interval for the mean known σ of the population

Confidence interval for When the population is normally the mean ̅ ~ (, ) distributed, the distribution of the mean is (known σ of √ also normal the population) μ A 100(1 − α)% confidence interval for when population variable X is normally distributed and is known. ̅ ̅( − , + )⁄ ⁄ 2 2 = quantile of the normal probability √ √ ⁄2 distribution, that can be found from statistical tables.

Confidence interval for the mean unknown σ of the population

Confidence interval for Given that σ is not know we need to use its ̅ − the mean estimator S. = The random variable t has the Student’s t (unknown σ of √ distribution with n − 1 degrees of freedom. the population) ̅ ̅( − , + ) A 100(1−α)% confidence interval for μ, ⁄ ⁄ 2 2 √ √ where t is the upper α/2 point of the α/2 Student’s t distribution with n − 1 degrees 2−̅( )∑ of freedom. √ =1 = Where −1.

Hypothesis testing

Hypothesis Testing To verify that the mean of a population is equal to a certain value μ, against the alternative hypothesis of a value different ̅ Test for the − 0 from it, if we know σ, we can use the test = mean (known statistics Z. If Z has values near 0 we variance) √ cannot reject H , else we refuse H (two 0 0 side test). : value of our null hypothesis 0.

Usually we do not know σ and we estimate Test for the ̅ − it through S. In this case the test statistics 0 mean = to be use is t. (unknown It has the Student’s t distribution with n − 1 √ variance) degrees of freedom if H0 is true.

A. Sample size determination using statistical formulae: the confidence interval approach

A. Sample size determination using statistical formulae: The confidence interval approach 2 2 ̅ − = From = 2 √ The average ̅ = − = we have N.B: width of the C.I. is the double of the error √ parameter We want to obtain a confidence interval ℎ = 2 ⋅ = 2 ⋅ √ for the mean with an error equal to e.

B. Simple linear regression

B. Simple linear regression We want to estimate If there is a deterministic relationship f between a Deterministic + Y = dependent variable Y and an independent 0 1 relationship variables X. In the most simple case (linear and k = 1).

Deterministic relationship

Deterministic + Y = dependent variable Y and an independent 0 1 relationship variables X. In the most simple case (linear and k = 1).

Stochastic regression model

The term ε is an assumption, that Stochastic incorporates the size of the approximation + Y = +ε regression 0 1 error. The introduction of such element model identifies the stochastic relationship.

Ordinary least squares (OLS)

This method minimizes the following Ordinary Least Squares 2 auxiliary function (we want to minimize the 2 ̂ ̂̂∑ ( ) ∑ ( )− = − − 0 1 distance between our Y and the real Y that Squares (OLS) =1 =1 line in our line, so the error) ̅)( ̅) The Slope:( =∑ − − ∑ =1 ̂̂ = = - numerator = numerator of covariance ̅) 2(∑ ∑− - denominator = deviance =1 =.

N.B: From now on small letter means that that value is The change of Y caused by the unitary estimated as difference from the mean change in X.

The intercept

The Intercept: ̂ ̂ ̂̅ ̅ = − the average value that you expect for your 0 1 dependent variables when you have X=0.

Gauss Markov theorem

Enables us to proof that our estimator are Gauss Markov the best one that we could find, as in they (guardare dimostrazioni appunti) are the estimators that have minimum Theorem variance.

OLS estimator distribution

The proof of the normality distribution of 2 now is needed, because we want to be ̂ ~ ( , ) sure that the model we have chosen is 1 1 2∑ OLS estimator correct. 2 distribution ̂̂ As is a weighted average of y and the y 2~ ( , ) 0 0 2∑ ̂ are normally distributed, is also normally distributed.

Residual variance estimator

- n-2 (the degrees of freedom of our linear regression model) where 2 correspond to 2 the number of parameter we’re estimating ̂ 2 )∑( − − ∑̂ 0 1 Residual variance 2 for our model. (In the general formula it = ̂ = =2 −2 −2 Variance we’ll be n-k, where k is the number of our parameter) Estimator ̂̂ = − (Residual) - The residual (the distance of each point from the line) can be used as an estimator of the error.

OLS estimators standard error

22 Once we have estimated the variance of = OLS ̂ 2∑1 the stochastic term in the regression Estimators model, we may substitute it in the OLS standard 2 estimator’s variance to obtain the standard ∑2 2 error = ( ) errors (s.e.) of estimates. ̂ 2∑0.

C. Simple linear regression inference

C. Simple linear regression inference We can standardize OLS estimator distribution: (remember that standardize Standardize a ̂ − means subtracting to the estimator the 1 1 ~(0,1) Gaussian mean and dividing by standard deviation) -2 2 √ ∕ variable > then we have a standardized distribution with 0 mean and st. dev=1.

The t-distribution

It is a ratio of 2 distribution: a normal distribution for beta, and a chi-squared The t- distribution (a distribution which arise distribution from a normal distribution squared) for the Where is a t distribution with ν= n−2 degrees of variance. freedom.

CI confidence interval limits

CI (confidence The CI for β at a (1−α)% can be calculated as: interval) limits : β1 = 0 -> we want to reject our null 0 hypothesis (that the parameter is =0), because it means that the parameter is we reject the null hypothesis if: Hypothesis significative for our model, and so we’ve ̂ Testing 1| | > ; − 2 done a good estimation. β1 = 0 ̂ 2 : β1 ≠ 0 1 (so rejecting is positive for the model, while accepting not).

Here we’re not testing the importance of ̂ − our parameter in our model, we’re are 1| | > ; − 2 β1 = c testing the similarity with a specific value ̂ 21 (so accept is positive).

Relationship and dependence

One must not mistake the Pearson coefficient with the regression coefficient β1 of a simple linear regression. They can be linked by the following formula.

N.B: The Pearson coefficient (r) measures the Relationship = = = linear bidirectional relationship between X 1 and and Y. Dependence The regression coefficient β1 measures the linear dependence of Y from X. Regression analysis is based of causal links.

Residual analysis

Ex post analysis checking situation of the TSS = RSS + ESS error. Remember that we’ve estimated the (Total sum of squares = Residual sum of squares + Residual error through the residual, which are the Explained sum of squares) distances, the differences between the real Analysis ̂ of your data, and the fitted of your 2 2 2∑ = ∑̂ + ∑̂ model.

̅)2 2∑(∑ = − TSS 2 ̂2 RSS )∑̂ = ∑( − 2 ̂∑(̂ ̅)2 212 ESS ∑̂ = − = Tells you the percentage of original 2 : coefficient = = 1− variability catches by your model, and so of linear how good your model is. determination 2() Larger the coefficient, better is the model = → 0 ≤ ≤ 1.

There is a relationship between the 2 coefficient of linear determination and the 2 2 ∑̂ ∑ ̂ (̂2 2()) = = = = = Pearson correlation coeff. 1 12 2 ∑ ∑ N.B: you have to look at beta, because the and r correlation coeff. can be also negative, and ( = 1 − ) if I squared it, it comes out always as positive. 2(̂ 2)− 1 1 12~2.

Chi-squared distribution

Chi-squared distribution As the square of a Normal distribution distribution ∑ 2. .: ~−22 The ratio of χ2 divided by the d.f.:-> ratio of the part of the variability catches by the model, divided by the part of the variability unexplained.

F distribution

2 -> higher the ratio, better is our model. (̂ 2)− 1 1 F distribution -> N.B: if you have to test the significance ~ ; − 212 ( 2) ∕ − of parameters, you have to perform multiple test for each variability. If you want to test the significance of the model as a whole, you’ve to use the F distribution.

A strong linear relationship between X and Y will determine a high valued test statistic, which supports the model, and reject the verify the null null hypothesis. Therefore: 2(̂ 2 hypothesis: ) ∕ 11 = ~ 1;−2 : β1 = 0 ( 2) ( 2)∑ ∕ − ∕ − : β1 ≠ 0 we want our hypothesis to be rejected, because =0 means that our model is not significant.

Analysis of variance (ANOVA table)

Analysis Of Sum of Mean F-statistic d.f. VAriance Squares Square (ANOVA table) 22 2 Estimated ∑̂ ∕ 11 ∑̂ ∑̂ 21 ∑̂ ∕ − 22 2 Residual ∑̂ ∑̂ n-2 −22 Total n-1 ∑

Forecast prevision

Often the estimated regression model is ̂ used to predict the expected Y value (i.e. Forecast ̂ ̂̂ = + ) which corresponds to a specific value of 0 1 (prevision) the X variable in this case indicated with Xκ.

Numerator: if the value you want to forecast is very ̅)2(1 −-> Standard distant from the mean of observed value, its value will (̂ √1). . = + + increase, and the variability will be very affected (and ̅) Error 2∑( − a confidence interval will be very large).

D. Multiple linear regression model

D. Multiple linear regression model The Model Standard notation: for each statistical unit = + + ⋯ + 1 1 2 2 i=1...n specification Y = Xβ +ε Y : (n × 1) vector of n dependent variable observations Matrix X : (n × k) matrix of k regressors with each notation n observations β : (k × 1) vector of k parameters ε : (n × 1) vector of n.

OLS estimator

The OLS estimator in multiple linear ̂ regression is the vector that minimize 2 ̃)2∑ = ∑( − → ̃ the function of as the minimal sum of OLS estimator squares of the residual, which is the =1 difference of real Y and fitted Y (presented ̃ ̂′ −1 ′( )= ≡ with its formula), where is the i-th row of the X matrix.

Variance of the OLS estimator

(̂) 2 ′ −1( )= Variance of the OLS estimator (̂ 2 ′ −1)[( ]) = Unbiased ′ 2 estimator of ̂ = − parameter 20 ≤ ≤ 12 = ≥0→ 2 never decreases when a new X variable is added to the model -> this can be a 2 − ∑2 = =1− ≤1 disadvantage when comparing different 2 ∑ models.

Adjusted

What is the net effect of adding a new 2 ⁄ ( )∑ − variable? 2 =1− = 2 ⁄( 1)∑ − - we lose a degree of freedom when a new Adjusted ( 1) X variable is added − 2(1 )=1− − 2 - The adjusted adjusts for the number ( )− of variables (k).

ANOVA table and F-test

  • Find the 95th or the 99th quantile of the SS d.f. Mean F-statistic Sq. distribution (−1),(−) ̂ ′ ′ −1 Model =
  • If F > one rejects ( ∕ − 1) (1−);(−1),(−) ′ 2 . .= = ( ∕ − )

ANOVA table F-test for the overall significance of the ′ = − Residual model. It shows if there is a linear 1 2 )= (1 − . . relationship between all of the X variables ′ = −1 Total 2= ∑ considered together and Y.

Single parameter

̂ −02 1 ~(0,1) if is known, under : To test the hypothesis if the individual 0 2 ′ −1 Single [( ]√ ) variable Xi has a significant effect on Y we parameter have to test: 2 2.

Anteprima
Vedrai una selezione di 20 pagine su 96
Statistics for business decision making + Formulario Pag. 1 Statistics for business decision making + Formulario Pag. 2
Anteprima di 20 pagg. su 96.
Scarica il documento per vederlo tutto.
Statistics for business decision making + Formulario Pag. 6
Anteprima di 20 pagg. su 96.
Scarica il documento per vederlo tutto.
Statistics for business decision making + Formulario Pag. 11
Anteprima di 20 pagg. su 96.
Scarica il documento per vederlo tutto.
Statistics for business decision making + Formulario Pag. 16
Anteprima di 20 pagg. su 96.
Scarica il documento per vederlo tutto.
Statistics for business decision making + Formulario Pag. 21
Anteprima di 20 pagg. su 96.
Scarica il documento per vederlo tutto.
Statistics for business decision making + Formulario Pag. 26
Anteprima di 20 pagg. su 96.
Scarica il documento per vederlo tutto.
Statistics for business decision making + Formulario Pag. 31
Anteprima di 20 pagg. su 96.
Scarica il documento per vederlo tutto.
Statistics for business decision making + Formulario Pag. 36
Anteprima di 20 pagg. su 96.
Scarica il documento per vederlo tutto.
Statistics for business decision making + Formulario Pag. 41
Anteprima di 20 pagg. su 96.
Scarica il documento per vederlo tutto.
Statistics for business decision making + Formulario Pag. 46
Anteprima di 20 pagg. su 96.
Scarica il documento per vederlo tutto.
Statistics for business decision making + Formulario Pag. 51
Anteprima di 20 pagg. su 96.
Scarica il documento per vederlo tutto.
Statistics for business decision making + Formulario Pag. 56
Anteprima di 20 pagg. su 96.
Scarica il documento per vederlo tutto.
Statistics for business decision making + Formulario Pag. 61
Anteprima di 20 pagg. su 96.
Scarica il documento per vederlo tutto.
Statistics for business decision making + Formulario Pag. 66
Anteprima di 20 pagg. su 96.
Scarica il documento per vederlo tutto.
Statistics for business decision making + Formulario Pag. 71
Anteprima di 20 pagg. su 96.
Scarica il documento per vederlo tutto.
Statistics for business decision making + Formulario Pag. 76
Anteprima di 20 pagg. su 96.
Scarica il documento per vederlo tutto.
Statistics for business decision making + Formulario Pag. 81
Anteprima di 20 pagg. su 96.
Scarica il documento per vederlo tutto.
Statistics for business decision making + Formulario Pag. 86
Anteprima di 20 pagg. su 96.
Scarica il documento per vederlo tutto.
Statistics for business decision making + Formulario Pag. 91
1 su 96
D/illustrazione/soddisfatti o rimborsati
Acquista con carta o PayPal
Scarica i documenti tutte le volte che vuoi
Dettagli
SSD
Scienze economiche e statistiche SECS-S/01 Statistica

I contenuti di questa pagina costituiscono rielaborazioni personali del Publisher Ari_Cora di informazioni apprese con la frequenza delle lezioni di Statistics for business decision making e studio autonomo di eventuali libri di riferimento in preparazione dell'esame finale o della tesi. Non devono intendersi come materiale ufficiale dell'università Università degli Studi di Siena o del prof Gagliardi Francesca.
Appunti correlati Invia appunti e guadagna

Domande e risposte

Hai bisogno di aiuto?
Chiedi alla community