This is my first python paper!¶
Section 1¶
import numpy as np
x = [5, 2, 3]
type(x)
# list
y = ["a", "b", "c"]
z = [True, False, False, True]
x[0]
5
::: {.callout-warning} Watch out! Python indexing starts at 0, not 1 as in R! :::
y[0] # 'a' <- first element
y[1] # 'b' <- second element
'b'
Lists: Printing, Concatenating, Coercing¶
print("The first value in y is '", y[0], "'.", sep="")
# The first value in y is 'a'.
"The first value in y is " + y[0] + "."
# 'The first value is y is a.'
str(x[0]) # '1' -- coerce a number to a string
The first value in y is 'a'.
'5'
A list can build a sequence with range, and can mix types freely:
u = list(range(0, 10, 2)) # [0, 2, 4, 6, 8]
v = [True, 0, "whatever", [1, 2]]
type(v[2]) # str
str
Lists: Multiple Assignment and Editing¶
t, u, v = 3, [4, 5], "hello"
tuv = [t, u, v]
tuv[2] = "good-bye"
tuv
[3, [4, 5], 'good-bye']
Important contrast with R: the * operator on a list does not do entrywise arithmetic --- it repeats the list!
x * 2
# [1, 2, 3, 1, 2, 3]
[5, 2, 3, 5, 2, 3]
This is one of several signs that lists are not built for numerical computing --- that's what NumPy arrays are for.
The Python Tuple¶
A tuple is like a list, but immutable --- it cannot be edited after creation. Built with parentheses:
salutations = ("hey", "hi", "ahoy", "sup")
salutations[0] # 'hey'
list(salutations) # convert to a list
['hey', 'hi', 'ahoy', 'sup']
- Trying to edit a tuple's entry raises an error
- Immutability makes tuples more memory-efficient
- We will mostly see tuples used as arguments to functions (e.g. specifying the shape of an array)
The Python Dictionary¶
A dictionary stores key--value pairs, accessed by key rather than by position:
stat_540 = {'nstudents' : 29,
'time' : '2:20 - 3:35 pm',
'days' : ['Tue', 'Thu'],
'instructor' : 'Dr. Huang'}
stat_540['time'] # '1:15 - 2:30 pm'
'2:20 - 3:35 pm'
Values can be edited by key:
stat_540['instructor'] = 'Dr. Ho'
For most of our statistics and data-science work, we will rely much more on NumPy arrays than on lists or dictionaries.
Outline¶
- Why Python? Getting oriented
- Basic Python objects: lists, tuples, dictionaries
- NumPy arrays: creating, indexing, slicing
- Array attributes, reshaping, and combining arrays
- Arithmetic and summary statistics on arrays
- Practice
Creating NumPy Arrays¶
NumPy arrays are more like R's vectors and matrices: entries must all be the same type.
import numpy as np
x = np.array([0, 1, 2])
type(x) # numpy.ndarray
numpy.ndarray
Building sequences, similarly to seq() in R:
seq = np.arange(-1, 1, 1/4) # [start, stop), step
seq2 = np.linspace(0, 1, 21) # 21 equally-spaced points on [0,1]
print(seq)
print(seq2)
[-1. -0.75 -0.5 -0.25 0. 0.25 0.5 0.75] [0. 0.05 0.1 0.15 0.2 0.25 0.3 0.35 0.4 0.45 0.5 0.55 0.6 0.65 0.7 0.75 0.8 0.85 0.9 0.95 1. ]
Mixed types get upcast (e.g. integers become floats) so the array stays a single type.
Creating Arrays from Scratch¶
np.zeros(10, dtype=int) # length-10 array of zeros
np.ones((3, 5)) # 3x5 array of ones
np.full((3, 7), 4) # 3x7 array, all entries = 4
np.eye(3) # 3x3 identity matrix
np.diag(np.ones(5)) # 5x5 diagonal / identity matrix
np.random.random((3, 3)) # Uniform(0,1) entries
np.random.normal(0, 1, (3, 3)) # Normal(0,1) entries
np.random.randint(0, 10, (3, 3)) # random integers in [0,10)
array([[4, 2, 2],
[7, 4, 3],
[9, 6, 1]])
np.random.normal(0, 1, (3, 3))
array([[ 0.61342943, -0.70669823, -0.51996023],
[ 0.29909795, -0.06506812, 1.0820678 ],
[-1.06165841, 0.95420008, -0.09446331]])
- Give
(rows, columns)as a tuple for shape - A matrix (2-d array) is made from a list of equal-length lists:
np.array([[1,2,3],[4,5,6]])
np.array([[1,2,3],[4,5,6]])
array([[1, 2, 3],
[4, 5, 6]])
Random Number Generators¶
The currently recommended way to draw random numbers is via a generator object:
rng = np.random.default_rng()
rng.choice(['apple', 'banana', 'cherry'], size=10)
array(['banana', 'apple', 'apple', 'banana', 'banana', 'cherry', 'apple',
'banana', 'banana', 'cherry'], dtype='<U6')
Randomly picks 10 items (with replacement by default).
index=rng.choice(np.arange(0, 10), size=10)
index
array([5, 6, 1, 1, 6, 7, 1, 6, 7, 3])
my_array=np.arange(0,10)
rng.shuffle(my_array)
print(my_array)
[1 6 3 4 2 8 5 0 9 7]
X = rng.poisson(lam=2, size=20)
X = rng.poisson(lam=2, size=(10, 3)) # tuple size -> 2-d array
U = rng.random((4, 3)) # Uniform(0,1)
B = rng.binomial(1, 1/2, (4, 3)) # Bernoulli(1/2)
print(B)
[[0 0 1] [1 0 0] [1 1 0] [1 1 0]]
This mirrors rpois(), runif(), rbinom() in R, but the distribution's parameters are accessed as methods of the generator object.
# Pass any integer as a seed
rng = np.random.default_rng(seed=42)
print(rng.random()) # This will output the exact same number every single time
0.7739560485559633
Exercise¶
(1) Generate 30 random samples from Normal(0,1)
(2) Randomly assign 15 samples into control and treatment group
n=15
sam = rng.normal(loc=0, scale=1, size=2*n) # rnorm(30)
sam[:5]
array([-0.14338303, -0.9611649 , -0.57378864, -0.643139 , -2.91023702])
group=np.repeat([0,1], 15)
group
array([0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 1, 1, 1, 1,
1, 1, 1, 1, 1, 1, 1, 1])
sgroup=rng.permutation(group)
sgroup
array([0, 1, 1, 1, 1, 0, 0, 0, 0, 1, 1, 1, 0, 1, 0, 0, 0, 0, 0, 0, 1, 1,
1, 0, 0, 1, 1, 1, 0, 1])
control=sam[sgroup==0]
trt=sam[sgroup==1]
control
array([-0.14338303, 1.75494901, 0.63006027, -0.42047308, 0.76775837,
1.26134475, -0.03232397, -0.96025364, 2.10660849, 1.56855716,
-0.01562377, -0.87418331, -0.33996295, -0.18621914, 1.34673547])
trt
array([-0.9611649 , -0.57378864, -0.643139 , -2.91023702, 0.42213754,
-1.97139064, -0.54171989, -1.155092 , 0.69490999, 1.53274065,
1.10735997, -0.01341163, 0.0112715 , -0.24035469, -1.21511317])
Accessing Array Entries¶
seq = np.arange(-1, 1, 1/4)
print(seq)
print(seq[0]) # first entry: -1.0
print(seq[2:4]) # entries with index 2, 3
print(seq[3:]) # from index 3 to the end
print(seq[:4]) # from the start up to (not incl.) index 4
print(seq[::3]) # every third entry
print(seq[-1]) # last entry (negative index counts from the end)
[-1. -0.75 -0.5 -0.25 0. 0.25 0.5 0.75] -1.0 [-0.5 -0.25] [-0.25 0. 0.25 0.5 0.75] [-1. -0.75 -0.5 -0.25] [-1. -0.25 0.5 ] 0.75
Key difference from R: In Python, a negative index counts entries from the end; it does not drop an entry as it does in R.
Indexing 2-D Arrays (Matrices)¶
M = np.array([[1, 2, 3], [4, 5, 6]])
print(M)
print(M[0, :]) # first row: [1 2 3]
print(M[:, 1]) # second column: [2 5]
print(M[1, 2]) # single entry: 6
M[1, 2] = 7 # replace an entry
M
[[1 2 3] [4 5 6]] [1 2 3] [2 5] 6
array([[1, 2, 3],
[4, 5, 7]])
De-selecting rows/columns uses ~ rather than a negative index:
print(M[~0, :]) # drop the first row
even = M % 2 == 0
print(even)
M[~even] # boolean mask: keep the odd entries
[4 5 7] [[False True False] [ True False False]]
array([1, 3, 5, 7])
Outline¶
- Why Python? Getting oriented
- Basic Python objects: lists, tuples, dictionaries
- NumPy arrays: creating, indexing, slicing
- Array attributes, reshaping, and combining arrays
- Arithmetic and summary statistics on arrays
- Practice
Array Attributes¶
Attributes describe an array's dimensions and storage type; access with a . (no parentheses):
A = np.random.randint(10, size=(3, 4, 5))
A.ndim # number of dimensions: 3
A.shape # size of each dimension: (3, 4, 5)
A.size # total number of entries: 60
A.dtype # data type, e.g. dtype('int64')
dtype('int64')
Every NumPy array has a single dtype: common ones are int64, float64, bool_.
Watch out! Assigning a float into an integer array silently truncates the value --- no warning is given!
A=np.random.randint(10, size=(3,4))
A
array([[1, 7, 3, 2],
[6, 9, 5, 8],
[7, 3, 8, 6]])
A[0,0]=2.2
A
array([[2, 7, 3, 2],
[6, 9, 5, 8],
[7, 3, 8, 6]])
A = A.astype(float)
A[0,0]=2.2
A
array([[2.2, 7. , 3. , 2. ],
[6. , 9. , 5. , 8. ],
[7. , 3. , 8. , 6. ]])
Reshaping Arrays¶
np.reshape(np.linspace(0, 1, 21), (3, 7))
array([[0. , 0.05, 0.1 , 0.15, 0.2 , 0.25, 0.3 ],
[0.35, 0.4 , 0.45, 0.5 , 0.55, 0.6 , 0.65],
[0.7 , 0.75, 0.8 , 0.85, 0.9 , 0.95, 1. ]])
np.reshape(np.linspace(0, 1, 21), (7, 3))
array([[0. , 0.05, 0.1 ],
[0.15, 0.2 , 0.25],
[0.3 , 0.35, 0.4 ],
[0.45, 0.5 , 0.55],
[0.6 , 0.65, 0.7 ],
[0.75, 0.8 , 0.85],
[0.9 , 0.95, 1. ]])
reshapefills the new array across rows first- Use
np.transpose(A)(orA.transpose()) to fill across columns instead - Arrays can have more than 2 dimensions, just as in R
np.transpose(np.reshape(np.linspace(0, 1, 21), (3, 7)))
array([[0. , 0.35, 0.7 ],
[0.05, 0.4 , 0.75],
[0.1 , 0.45, 0.8 ],
[0.15, 0.5 , 0.85],
[0.2 , 0.55, 0.9 ],
[0.25, 0.6 , 0.95],
[0.3 , 0.65, 1. ]])
Methods: Applying Functions to Objects¶
Many NumPy functions can also be called as a method attached to the object with a .:
B = np.array([[1, 2], [3, 4]])
B
array([[1, 2],
[3, 4]])
np.transpose(B) # function form
B.transpose() # method form -- same result
array([[1, 3],
[2, 4]])
np.sum(B) # function form
B.sum() # method form -- same result
np.int64(10)
This "object.method()" pattern has no direct analogue in base R and shows up constantly in Python.
Aliases vs. Copies¶
A subset of an array is an alias, not a new array --- editing it edits the original!
A = np.reshape(np.arange(24), (4, 6))
A
array([[ 0, 1, 2, 3, 4, 5],
[ 6, 7, 8, 9, 10, 11],
[12, 13, 14, 15, 16, 17],
[18, 19, 20, 21, 22, 23]])
A0 = A[:2, :2]
A0[1, 1] = -99
A
# A has ALSO changed!
array([[ 0, 1, 2, 3, 4, 5],
[ 6, -99, 8, 9, 10, 11],
[ 12, 13, 14, 15, 16, 17],
[ 18, 19, 20, 21, 22, 23]])
To get an independent copy, append .copy():
A0_copy = A[:2, :2].copy()
A0_copy[0, 0] = -99 # does NOT affect A
A
array([[ 0, 1, 2, 3, 4, 5],
[ 6, -99, 8, 9, 10, 11],
[ 12, 13, 14, 15, 16, 17],
[ 18, 19, 20, 21, 22, 23]])
Subsetting with a boolean mask¶
A[A < 5] = 0 # set to zero all entries in A which are less than 0.5
A
array([[ 0, 0, 0, 0, 0, 5],
[ 6, 0, 8, 9, 10, 11],
[12, 13, 14, 15, 16, 17],
[18, 19, 20, 21, 22, 23]])
Combining and Splitting Arrays¶
x = np.array([1, 2, 3])
y = np.array([98, 99, 100])
np.concatenate([x, y]) # [1 2 3 98 99 100]
array([ 1, 2, 3, 98, 99, 100])
u = rng.binomial(1,1/2,(4,3)) # 4 by 3
v = rng.random((2,3)) # 2 by 3
print(u)
print(v)
print(u[2:])
[[1 0 1] [1 1 1] [0 1 0] [1 0 0]] [[0.60372623 0.37617457 0.70187553] [0.22796474 0.88346417 0.17529252]] [[0 1 0] [1 0 0]]
np.vstack([u,v]) # stack rows (matching # cols)
array([[1. , 0. , 1. ],
[1. , 1. , 1. ],
[0. , 1. , 0. ],
[1. , 0. , 0. ],
[0.60372623, 0.37617457, 0.70187553],
[0.22796474, 0.88346417, 0.17529252]])
np.hstack([u[2:], v]) # stack columns (matching # rows)
array([[0. , 1. , 0. , 0.60372623, 0.37617457,
0.70187553],
[1. , 0. , 0. , 0.22796474, 0.88346417,
0.17529252]])
Splitting is the reverse operation:
x1, x2, x3 = np.split(np.arange(12), [3, 6])
print(x1)
print(x2)
print(x3)
[0 1 2] [3 4 5] [ 6 7 8 9 10 11]
Outline¶
- Why Python? Getting oriented
- Basic Python objects: lists, tuples, dictionaries
- NumPy arrays: creating, indexing, slicing
- Array attributes, reshaping, and combining arrays
- Arithmetic and summary statistics on arrays
- Practice
Arithmetic on NumPy Arrays¶
The operators + - * / and ** (exponent), % (modulus) work entrywise, as in R:
a = np.array([3, 4, 5])
c = 2
a + c # [5 6 7]
array([5, 6, 7])
a ** c # [9 16 25] ('^' is NOT exponentiation in Python!)
array([ 9, 16, 25])
a // c # floor divide: [1 2 2]
array([1, 2, 2])
No recycling! Unlike R, NumPy will not silently recycle a shorter
array to match a longer one --- mismatched shapes raise an error.
Common math functions live in NumPy: np.exp, np.log, np.sin, np.arctan, ...
Summary Statistics on Arrays¶
D = rng.random((20, 3))
print(D)
np.sum(D) # or D.sum()
np.min(D) # or D.min()
np.max(D) # or D.max()
[[0.62695242 0.90651526 0.07511508] [0.8588796 0.44915439 0.0346995 ] [0.20957312 0.12097211 0.89711882] [0.66346972 0.24289914 0.96323799] [0.98911702 0.65459397 0.91216912] [0.00739 0.31258323 0.11407253] [0.32975446 0.63630544 0.0810894 ] [0.15217003 0.93215753 0.88304247] [0.56933459 0.45617789 0.31308781] [0.29629497 0.79257319 0.27125197] [0.58167581 0.17638694 0.84098462] [0.75681068 0.08562706 0.42929889] [0.8166283 0.23701507 0.74515563] [0.09847735 0.70850438 0.39360128] [0.51253678 0.01925071 0.25216621] [0.17066232 0.40323255 0.63050813] [0.32971842 0.7170679 0.71863739] [0.44845096 0.1447243 0.30643352] [0.12308156 0.88863459 0.39452356] [0.68892699 0.20626377 0.06670593]]
0.989117016557217
D.sum(axis=0)
array([9.22990509, 9.09063942, 9.32289987])
D = rng.random((20, 3))
np.sum(D) # or D.sum()
np.min(D) # or D.min()
np.max(D) # or D.max()
D.sum(axis=0) # column sums
D.max(axis=1) # row maxima
array([0.61810602, 0.81397754, 0.5088196 , 0.94443434, 0.78419807,
0.99668802, 0.79078276, 0.92715229, 0.8358856 , 0.69289495,
0.8427604 , 0.84561569, 0.96577913, 0.75940195, 0.56301702,
0.7367655 , 0.52794625, 0.55931858, 0.67722432, 0.88355547])
axis=0operates down columns;axis=1operates across rowsnp.sumis much faster than Python's built-insumon large arrays
Missing Values¶
Create a missing value with None; most summary functions have a nan-aware version:
size=D.shape
size
(20, 3)
rng.binomial(1, 0.1, size=D.shape)
array([[1, 0, 0],
[0, 0, 0],
[0, 0, 0],
[0, 0, 0],
[0, 0, 0],
[0, 0, 0],
[0, 0, 0],
[1, 0, 0],
[0, 0, 0],
[0, 0, 0],
[0, 0, 0],
[0, 0, 0],
[0, 0, 1],
[0, 0, 0],
[0, 1, 0],
[0, 0, 0],
[0, 0, 0],
[0, 0, 0],
[0, 0, 0],
[0, 0, 0]])
D[rng.binomial(1, 0.1, size=D.shape) == 1] = None
print(D)
np.nanmax(D) # ignores missing values
np.nanmin(D, axis=0) # column minima, ignoring NaNs
[[0.62695242 0.90651526 0.07511508] [0.8588796 0.44915439 0.0346995 ] [0.20957312 0.12097211 0.89711882] [0.66346972 0.24289914 0.96323799] [0.98911702 0.65459397 0.91216912] [0.00739 0.31258323 0.11407253] [0.32975446 0.63630544 0.0810894 ] [0.15217003 0.93215753 0.88304247] [ nan 0.45617789 0.31308781] [0.29629497 0.79257319 0.27125197] [0.58167581 0.17638694 0.84098462] [0.75681068 0.08562706 0.42929889] [ nan 0.23701507 0.74515563] [0.09847735 0.70850438 0.39360128] [0.51253678 0.01925071 0.25216621] [0.17066232 0.40323255 0.63050813] [0.32971842 0.7170679 0.71863739] [0.44845096 0.1447243 0.30643352] [0.12308156 0.88863459 0.39452356] [0.68892699 0.20626377 nan]]
array([0.00739 , 0.01925071, 0.0346995 ])
Compare: np.sum $\to$ np.nansum, np.mean $\to$ np.nanmean, etc. --- the same pattern as na.rm = TRUE in R, just with a different function name instead of an argument.
Outline¶
- Why Python? Getting oriented
- Basic Python objects: lists, tuples, dictionaries
- NumPy arrays: creating, indexing, slicing
- Array attributes, reshaping, and combining arrays
- Arithmetic and summary statistics on arrays
- Practice
Practice: Write Code¶
Write code to create the list
[0, 3, 6, 9, 12, 0, 3, 6, 9, 12].Create the NumPy array whose $i$th row (starting at $i=0$) is filled with the value $i$, repeated 4 times, for $i = 0, \dots, 7$.
Write code to build the $(n-1) \times n$ "successive difference" matrix with $-1$ on the main diagonal and $1$ on the diagonal just above it, and zeros elsewhere, for any $n$.
Simulate 10,000 rolls of a pair of six-sided dice with
rng.integers(low=1, high=7, size=(10000,2)). Find the proportion of rolls whose sum is odd.
import numpy as np
x=list(np.arange(0,14, 3))*2
x
[0, 3, 6, 9, 12, 0, 3, 6, 9, 12]
A=np.zeros((8,4), dtype=int)
for i in range(8):
A[i, :] = i
print(A)
[[0 0 0 0] [1 1 1 1] [2 2 2 2] [3 3 3 3] [4 4 4 4] [5 5 5 5] [6 6 6 6] [7 7 7 7]]
A = np.array([[i] * 4 for i in range(8)])
A = np.arange(8).repeat(4).reshape(8, 4)
print(A)
[[0 0 0 0] [1 1 1 1] [2 2 2 2] [3 3 3 3] [4 4 4 4] [5 5 5 5] [6 6 6 6] [7 7 7 7]]
n=4
D=np.zeros((n-1, n))
D
array([[0., 0., 0., 0.],
[0., 0., 0., 0.],
[0., 0., 0., 0.]])
np.fill_diagonal(D, -1)
D
np.fill_diagonal(D[:, 1:], 1)
D
array([[-1., 1., 0., 0.],
[ 0., -1., 1., 0.],
[ 0., 0., -1., 1.]])
def diff_matrix(n):
D = np.zeros((n - 1, n))
np.fill_diagonal(D, -1) # -1 on main diagonal
np.fill_diagonal(D[:, 1:], 1) # +1 on diagonal just above
return D
print(diff_matrix(5))
[[-1. 1. 0. 0. 0.] [ 0. -1. 1. 0. 0.] [ 0. 0. -1. 1. 0.] [ 0. 0. 0. -1. 1.]]
rolls = rng.integers(low=1, high=7, size=(10000, 2)) # shape (10000, 2)
rolls.shape
# sum each pair
(10000, 2)
totals = rolls.sum(axis=1)
totals
prop_odd = np.mean(totals % 2 == 1)
print(f"Proportion of odd sums: {prop_odd:.4f}")
Proportion of odd sums: 0.5041
Practice: Read Code¶
Predict the output of each code chunk before running it.
# (1)
ch = "Why hello."
ch[2:]
# (2)
x = list(range(0, 15, 3))
x * 2
# (3)
a, b, c = [4, 5], ['cat', 'cow'], [True, False]
d = [a, b, c]
d[1][1]
'cow'
ch = "Why hello."
ch[2:]
'y hello.'
x = list(range(0, 15, 3))
x * 2
[0, 3, 6, 9, 12, 0, 3, 6, 9, 12]