Java正则表达式非贪婪不起作用

发布于 2025-01-04 21:39:13 字数 3310 浏览 5 评论 0原文

我正在尝试从以下 HTML 源代码中提取包含字符串“Active”的项目。请注意,活动项目可能会先出现。

这是我正在使用的正则表达式(尝试非贪婪):

<TR class=\\w*?Item>.*?Active.*?</TD></TR>

with Pattern.CASE_INSENSITIVE|图案.DOTALL | Pattern.MULTILINE

它提取整个源代码而不是第二个:...。

HTML 源代码:

<TR class=Item>
<TD style="WIDTH: 100px"><A id=ctl00_BodyContents_gvServers_ctl02_HyperLink1 onclick=javascript:turnColor(this); href="AddEditVirtualServer.aspx?ServerId=16733" target=_blank>server01</A> </TD>
<TD></TD>
<TD style="WIDTH: 100px"><A id=ctl00_BodyContents_gvServers_ctl02_HyperLink2 onclick=javascript:turnColor(this); href="AddEditVirtualServer.aspx?ServerId=16733" target=_blank>07D8F15</A> </TD>
<TD style="WIDTH: 150px">IBM 8204-E8A</TD>
<TD style="WIDTH: 150px"><SPAN>AIX - 5.3.0.0 - ML 12</SPAN></TD>
<TD style="WIDTH: 100px">PowerPC_POWER6</TD>
<TD style="WIDTH: 75px">1</TD>
<TD style="WIDTH: 100px">C1-D-G13</TD>
<TD style="WIDTH: 100px"><SPAN id=ctl00_BodyContents_gvServers_ctl02_lbLevel1SupportGroup>UNIX TEAM 1&amp;CORP</SPAN> </TD>
<TD style="WIDTH: 100px"><SPAN id=ctl00_BodyContents_gvServers_ctl02_lbLevel2SupportGroup>UNIX TEAM 1&amp;CORP</SPAN> </TD>
<TD style="WIDTH: 100px"><SPAN id=ctl00_BodyContents_gvServers_ctl02_lbLevel3SupportGroup>UNIX TEAM 1&amp;CORP</SPAN> </TD>
<TD style="WIDTH: 100px">2011-04-15</TD>
<TD style="WIDTH: 100px">Cool Down </TD>
<TD style="WIDTH: 100px"><A id=ctl00_BodyContents_gvServers_ctl02_btnDeleteServer disabled>Delete</A> </TD></TR>
<TR class=AlternateItem>
<TD style="WIDTH: 100px"><A id=ctl00_BodyContents_gvServers_ctl03_HyperLink1 onclick=javascript:turnColor(this); href="AddEditVirtualServer.aspx?ServerId=19631" target=_blank>server01</A> </TD>
<TD></TD>
<TD style="WIDTH: 100px"><A id=ctl00_BodyContents_gvServers_ctl03_HyperLink2 onclick=javascript:turnColor(this); href="AddEditVirtualServer.aspx?ServerId=19631" target=_blank>105ABCD</A> </TD>
<TD style="WIDTH: 150px">IBM Power 770</TD>
<TD style="WIDTH: 150px"><SPAN>AIX - 5.3.0.0 - TL 12 SP 01</SPAN></TD>
<TD style="WIDTH: 100px">PowerPC_POWER7</TD>
<TD style="WIDTH: 75px">1</TD>
<TD style="WIDTH: 100px">C1-O-G11</TD>
<TD style="WIDTH: 100px"><SPAN id=ctl00_BodyContents_gvServers_ctl03_lbLevel1SupportGroup>UNIX TEAM 1&amp;CORP</SPAN> </TD>
<TD style="WIDTH: 100px"><SPAN id=ctl00_BodyContents_gvServers_ctl03_lbLevel2SupportGroup>UNIX TEAM 1&amp;CORP</SPAN> </TD>
<TD style="WIDTH: 100px"><SPAN id=ctl00_BodyContents_gvServers_ctl03_lbLevel3SupportGroup>UNIX TEAM 1&amp;CORP</SPAN> </TD>
<TD style="WIDTH: 100px">2012-02-09</TD>
<TD style="WIDTH: 100px">Active </TD>
<TD style="WIDTH: 100px"><A id=ctl00_BodyContents_gvServers_ctl03_btnDeleteServer disabled>Delete</A> </TD></TR>

我们将不胜感激您的帮助!

I am trying to extract the Item that contain String "Active" from the following HTML source code. Please be noted the Active item may go first.

Here is the REGEX I am using(Trying to be non-greedy):

<TR class=\\w*?Item>.*?Active.*?</TD></TR>

with Pattern.CASE_INSENSITIVE| Pattern.DOTALL | Pattern.MULTILINE

It extract the whole source code instead of the second one: ... .

HTML Source code:

<TR class=Item>
<TD style="WIDTH: 100px"><A id=ctl00_BodyContents_gvServers_ctl02_HyperLink1 onclick=javascript:turnColor(this); href="AddEditVirtualServer.aspx?ServerId=16733" target=_blank>server01</A> </TD>
<TD></TD>
<TD style="WIDTH: 100px"><A id=ctl00_BodyContents_gvServers_ctl02_HyperLink2 onclick=javascript:turnColor(this); href="AddEditVirtualServer.aspx?ServerId=16733" target=_blank>07D8F15</A> </TD>
<TD style="WIDTH: 150px">IBM 8204-E8A</TD>
<TD style="WIDTH: 150px"><SPAN>AIX - 5.3.0.0 - ML 12</SPAN></TD>
<TD style="WIDTH: 100px">PowerPC_POWER6</TD>
<TD style="WIDTH: 75px">1</TD>
<TD style="WIDTH: 100px">C1-D-G13</TD>
<TD style="WIDTH: 100px"><SPAN id=ctl00_BodyContents_gvServers_ctl02_lbLevel1SupportGroup>UNIX TEAM 1&CORP</SPAN> </TD>
<TD style="WIDTH: 100px"><SPAN id=ctl00_BodyContents_gvServers_ctl02_lbLevel2SupportGroup>UNIX TEAM 1&CORP</SPAN> </TD>
<TD style="WIDTH: 100px"><SPAN id=ctl00_BodyContents_gvServers_ctl02_lbLevel3SupportGroup>UNIX TEAM 1&CORP</SPAN> </TD>
<TD style="WIDTH: 100px">2011-04-15</TD>
<TD style="WIDTH: 100px">Cool Down </TD>
<TD style="WIDTH: 100px"><A id=ctl00_BodyContents_gvServers_ctl02_btnDeleteServer disabled>Delete</A> </TD></TR>
<TR class=AlternateItem>
<TD style="WIDTH: 100px"><A id=ctl00_BodyContents_gvServers_ctl03_HyperLink1 onclick=javascript:turnColor(this); href="AddEditVirtualServer.aspx?ServerId=19631" target=_blank>server01</A> </TD>
<TD></TD>
<TD style="WIDTH: 100px"><A id=ctl00_BodyContents_gvServers_ctl03_HyperLink2 onclick=javascript:turnColor(this); href="AddEditVirtualServer.aspx?ServerId=19631" target=_blank>105ABCD</A> </TD>
<TD style="WIDTH: 150px">IBM Power 770</TD>
<TD style="WIDTH: 150px"><SPAN>AIX - 5.3.0.0 - TL 12 SP 01</SPAN></TD>
<TD style="WIDTH: 100px">PowerPC_POWER7</TD>
<TD style="WIDTH: 75px">1</TD>
<TD style="WIDTH: 100px">C1-O-G11</TD>
<TD style="WIDTH: 100px"><SPAN id=ctl00_BodyContents_gvServers_ctl03_lbLevel1SupportGroup>UNIX TEAM 1&CORP</SPAN> </TD>
<TD style="WIDTH: 100px"><SPAN id=ctl00_BodyContents_gvServers_ctl03_lbLevel2SupportGroup>UNIX TEAM 1&CORP</SPAN> </TD>
<TD style="WIDTH: 100px"><SPAN id=ctl00_BodyContents_gvServers_ctl03_lbLevel3SupportGroup>UNIX TEAM 1&CORP</SPAN> </TD>
<TD style="WIDTH: 100px">2012-02-09</TD>
<TD style="WIDTH: 100px">Active </TD>
<TD style="WIDTH: 100px"><A id=ctl00_BodyContents_gvServers_ctl03_btnDeleteServer disabled>Delete</A> </TD></TR>

Your help will be appreciated!

如果你对这篇内容有疑问,欢迎到本站社区发帖提问 参与讨论,获取更多帮助,或者扫码二维码加入 Web 技术交流群。

扫码二维码加入Web技术交流群

发布评论

需要 登录 才能够评论, 你可以免费 注册 一个本站的账号。

评论(2

不必在意 2025-01-11 21:39:13

*? 或多次。您可能需要

*? is zero or more times. You probably want <TR class=\\w+?Item>

余厌 2025-01-11 21:39:13

您可以通过 Pattern p = Pattern.compile("....", Pattern.MULTILINE); 启用多行模式
您可能必须将模式中的 更改为 .*?

You can enable multi-line mode by Pattern p = Pattern.compile("....", Pattern.MULTILINE);
And you probably have to change </TD></TR> to </TD>.*?</TR> in your pattern.

~没有更多了~
我们使用 Cookies 和其他技术来定制您的体验包括您的登录状态等。通过阅读我们的 隐私政策 了解更多相关信息。 单击 接受 或继续使用网站,即表示您同意使用 Cookies 和您的相关数据。
原文