基于另一个数据集获取数据集的子集
假设我有一个数据集,即 dat1
ID block plot SPID TotHeight
1 1 1 4 44.5
2 1 1 4 51
3 1 1 4 28.7
4 1 1 4 24.5
5 1 1 4 27.3
6 1 1 4 20
17 1 10 1 44.5
19 1 10 1 51
1 1 11 21 28.7
2 1 11 21 24.5
3 1 11 21 27.3
4 1 11 21 20
5 1 11 21 12.88666667
6 1 11 21 7.235238095
7 1 11 21 1.583809524
然后我有另一个大数据集,即 dat2:
ID block plot SPID Species TotHeight
1 1 1 4 BENI 72
2 1 1 4 BENI 55
3 1 1 4 BENI 51
4 1 1 4 BENI 47
5 1 1 4 BENI 49
6 1 1 4 BENI 34
7 1 1 4 BENI .
8 1 1 4 BENI 51
9 1 1 4 BENI 66
10 1 1 4 BENI 40
11 1 1 4 BENI 24
12 1 1 4 BENI 62
13 1 1 4 BENI 34
14 1 1 4 BENI 49
15 1 1 4 BENI 57
16 1 1 4 BENI 22
17 1 1 4 BENI 76
18 1 1 4 BENI 56
19 1 1 4 BENI 55
20 1 1 4 BENI 29
21 1 1 4 BENI 24
22 1 1 4 BENI 18
23 1 1 4 BENI 65
24 1 1 4 BENI 55
25 1 1 4 BENI 63
26 1 1 4 BENI 57
27 1 1 4 BENI 57
28 1 1 4 BENI 57
29 1 1 4 BENI 45
30 1 1 4 BENI 83
31 1 1 4 BENI 37
32 1 1 4 BENI 56
33 1 1 4 BENI 65
34 1 1 4 BENI 75
35 1 1 4 BENI 51
36 1 1 4 BENI .
1 1 2 16 PRSE 141
2 1 2 16 PRSE 192
3 1 2 16 PRSE .
4 1 2 16 PRSE 197
5 1 2 16 PRSE 172
6 1 2 16 PRSE 143
7 1 2 16 PRSE 141
8 1 2 16 PRSE 155
9 1 2 16 PRSE 167
10 1 2 16 PRSE 155
11 1 2 16 PRSE 175
12 1 2 16 PRSE 190
13 1 2 16 PRSE 148
14 1 2 16 PRSE 180
15 1 2 16 PRSE .
我的问题是如何从 dat2 中获取数据子集,其中 ID、块和图与 dat1 中的数据相匹配?如何从 dat2 中获取 ID、块和图与 dat1 中的数据不匹配的数据子集?
suppose I have one dataset, which is dat1
ID block plot SPID TotHeight
1 1 1 4 44.5
2 1 1 4 51
3 1 1 4 28.7
4 1 1 4 24.5
5 1 1 4 27.3
6 1 1 4 20
17 1 10 1 44.5
19 1 10 1 51
1 1 11 21 28.7
2 1 11 21 24.5
3 1 11 21 27.3
4 1 11 21 20
5 1 11 21 12.88666667
6 1 11 21 7.235238095
7 1 11 21 1.583809524
Then I have another big dataset, which is dat2:
ID block plot SPID Species TotHeight
1 1 1 4 BENI 72
2 1 1 4 BENI 55
3 1 1 4 BENI 51
4 1 1 4 BENI 47
5 1 1 4 BENI 49
6 1 1 4 BENI 34
7 1 1 4 BENI .
8 1 1 4 BENI 51
9 1 1 4 BENI 66
10 1 1 4 BENI 40
11 1 1 4 BENI 24
12 1 1 4 BENI 62
13 1 1 4 BENI 34
14 1 1 4 BENI 49
15 1 1 4 BENI 57
16 1 1 4 BENI 22
17 1 1 4 BENI 76
18 1 1 4 BENI 56
19 1 1 4 BENI 55
20 1 1 4 BENI 29
21 1 1 4 BENI 24
22 1 1 4 BENI 18
23 1 1 4 BENI 65
24 1 1 4 BENI 55
25 1 1 4 BENI 63
26 1 1 4 BENI 57
27 1 1 4 BENI 57
28 1 1 4 BENI 57
29 1 1 4 BENI 45
30 1 1 4 BENI 83
31 1 1 4 BENI 37
32 1 1 4 BENI 56
33 1 1 4 BENI 65
34 1 1 4 BENI 75
35 1 1 4 BENI 51
36 1 1 4 BENI .
1 1 2 16 PRSE 141
2 1 2 16 PRSE 192
3 1 2 16 PRSE .
4 1 2 16 PRSE 197
5 1 2 16 PRSE 172
6 1 2 16 PRSE 143
7 1 2 16 PRSE 141
8 1 2 16 PRSE 155
9 1 2 16 PRSE 167
10 1 2 16 PRSE 155
11 1 2 16 PRSE 175
12 1 2 16 PRSE 190
13 1 2 16 PRSE 148
14 1 2 16 PRSE 180
15 1 2 16 PRSE .
My question is how can I take a subset of data from dat2 in which ID, block and plot match those in dat1? And how can I get a subset of data from dat2 in which ID, block and plot do not match those in dat1?
如果你对这篇内容有疑问,欢迎到本站社区发帖提问 参与讨论,获取更多帮助,或者扫码二维码加入 Web 技术交流群。
绑定邮箱获取回复消息
由于您还没有绑定你的真实邮箱,如果其他用户或者作者回复了您的评论,将不能在第一时间通知您!
发布评论
评论(1)
听起来您想要对列 ID、块和图进行
合并
。假设您的数据以名为
dat1
和dat2
的形式读取,这应该就是您想要的:这实际上是在感兴趣的三列上以 SQL 术语执行内部联接。如果您也感兴趣,请阅读合并左、右和外连接可能性的帮助页面。
完全可重现的要点此处
要获取未从
dat2
合并的行,这个黑客有效。可能有一种更有效的方法可以做到这一点,但就是这样。首先,添加参数all.y = TRUE
。这指定了一个右连接,它将返回来自dat2
中未合并的行。然后我们可以对其进行子集化,因为知道未合并的行将返回 NA:It sounds like you want a
merge
on the columns ID, block, and plot.Assuming your data are read in named
dat1
anddat2
, this should be what you want:This essentially performs an inner join in SQL terms on the three columns of interest. Read the help page for merge for left, right, and outer join possibilities if that is also of interest.
Fully reproducible gist here
To get the rows that didn't merge from
dat2
, this hack works. There is probably a more efficient way to do this, but here it is. First, add the parameterall.y = TRUE
. This specifies a right join which will return rows fromdat2
which did not merge. Then we can subset on that knowing that the rows that didn't merge will returnNA
's: