Arguing about marketing ploys can be fun indeed, but completely impractical (unless you have access to a board-member of some frontier lab who's suddenly ready to get candid in public about this). Thus, that's not what I'm arguing about.
The more down-to-earth questions that interest me are:
1. Is there utility? Does it make sense to use those tools to find vulns in the code or it is a completely useless exercise?
2. How cost/benefit analysis changes this? Is it prohibitively expensive, or laying hands on these tools is impossible in practice, or there are some other obstacles/costs that make this non-viable?
My (speculative, using implicit evidence) argument was that the answers to these questions are "1. yes, there's one" and "2. not by much".
If from your POV the perceived utility is 0 because the real capability is so far below the advertised one, then it'd be a position, of course. But then there's a counter-argument of "how do you know if you haven't tried yet?".
Your own mention of "impressive findings" among "mundane observations, and garbage" tells me that it might be not as straightforward as this.
That's the questions I was looking to get answers for from you.
P.S.: Maybe I'm too early to the party. A year ago there were still a lot of skeptics (inc. in my circle) who were saying that agents will never be able to write proper code. Now most of them don't even read most of the code that an agent generates, except for the most critical paths. Maybe we'll just have to wait for the Chinese to release something that will force US labs to show their frontier to the public so everyone could try and see. If it all proves a failure, oddly enough I'll be among those celebrating this fact.
The more down-to-earth questions that interest me are:
1. Is there utility? Does it make sense to use those tools to find vulns in the code or it is a completely useless exercise?
2. How cost/benefit analysis changes this? Is it prohibitively expensive, or laying hands on these tools is impossible in practice, or there are some other obstacles/costs that make this non-viable?
My (speculative, using implicit evidence) argument was that the answers to these questions are "1. yes, there's one" and "2. not by much".
If from your POV the perceived utility is 0 because the real capability is so far below the advertised one, then it'd be a position, of course. But then there's a counter-argument of "how do you know if you haven't tried yet?".
Your own mention of "impressive findings" among "mundane observations, and garbage" tells me that it might be not as straightforward as this.
That's the questions I was looking to get answers for from you.
P.S.: Maybe I'm too early to the party. A year ago there were still a lot of skeptics (inc. in my circle) who were saying that agents will never be able to write proper code. Now most of them don't even read most of the code that an agent generates, except for the most critical paths. Maybe we'll just have to wait for the Chinese to release something that will force US labs to show their frontier to the public so everyone could try and see. If it all proves a failure, oddly enough I'll be among those celebrating this fact.
"