old-gregg wrote:
Yes, the AF points are on the sensor, but we're talking about subject recognition which is a separate step. [1]
For the AI model to find an eye, it needs to analyze the image. The camera doesn't feed the entire image into the AI chip, it's way too big. Instead, it crops it to AF area, and then downsamples. After downsampling, it becomes a very, very small. If the subject you're trying to detect occupies a small % of this cropped image, it won't be detectable after downsampling.
[1] With subject recognition enabled, focusing happens twice: first, the camera focuses optimizing for the entire AF area, then it runs subject recognition, and if it finds a subject, it fine-tunes the focus optimizing for what's found....Show more →
Also the readout speed is improved in crop-mode, because the sensor can be made to skip the top and bottom lines (and possibly left/right as well).
Something else to know about the AF points, there are 1 million+ such physical points and then they're grouped into logical points that perform the average of the physical points. That's why sometimes the number of AF points in APS-C mode doesn't necessarily need to be lower than in FF mode, because they're logical points.
Those points then provide a very small image like 33x23 (= 759) that can also be used for tracking subjects at a macro level. If the subject is smaller then the camera ideally would perform a ROI (region of interest) analysis covering one or a few of those macro points.
arbitrage wrote:
It is very much true for all cameras that have the "AI chip".
It is very easy to replicate over and over again. Find a bird far enough away that the camera doesn't recognize it as a bird....swap into APS-C mode and it will recognize it unless it is still too small.
Cameras before the AI chip actually performed the opposite. They would recognize birds in full frame mode and then have trouble with the same bird in APS-C.
Canon MILCs also perform the same way as the newer Sonys and again easy to replicate over and over again.
old-gregg wrote:
Jack, I am not the OP but I explained that in my previous comment above. The key error in your reasoning is that nothing is "magnified in the finder". Quite the opposite: the full 66MP image gets downsampled for the finder. But it also gets downsampled for the AF model, **massively** downsampled. So it is indeed quite helpful to have the bird fill up maximum % of the AF area, this is why tricks like shrinking the AF area (which is essentially what the OP is doing by switching to APS-C) help with recognition.
Neural network based subject recognition doesn't really work the way you seem to think it does.
Typically the first layer of the network will be fed the full image and from that it will extract bounding boxes for the detected objects e.g. body, head.
The word 'fed' is a little misleading here noting that the latest Sony SOC is a unified architecture so the neural engine probably has direct access to the shared memory storing the image and can scan the entire image directly. Typically these neural engines are optimised for massively parallel vector operations like scanning large images.
These bounding boxes will be fed to the next layer in the neural network for more refined extraction of features e.g. face, eyes, etc. which is likely then also used for AF calculations, movement prediction etc.
My experience with the Apple Vision APIs is that happens in realtime on every frame in the live video feed and you can get back the coordinates of the subject and any of the detected features (arms, legs, eyes, ears) so you can draw tracking objects like annotations in realtime on the screen. Bear in mind that in computer time 60 fps is quite slow.
I am pretty sure that's more or less how most of the modern mirrorless camera systems work in order to draw the subject detection boxes on screen and to determine the AF point(s), AF action required and AI based white balance adjustment.
hoodlum90 wrote:
Jan posted the 600mm review. He is planning another video very soon, comparing the 600, 400, 100-400, including TCs.
I was expecting much longer video from him since it came so late. Nothing new I guess. As a dual system user, I want to see comparison with Nikon counterparts. Sony need to improve OSS. So 400 might be same.
So far from all reviews and samples files, I think the 600 is equal or slightly better than Nikon version. 400 is visibly better than Nikon's 400 and adding 1.4x make it close to 600, but 600 is the way to go if that's the focal length you are after. I'm positive that 600 will beat 300+2X in IQ and AF.
Jemini wrote:
I was expecting much longer video from him since it came so late. Nothing new I guess. As a dual system user, I want to see comparison with Nikon counterparts. Sony need to improve OSS. So 400 might be same.
So far from all reviews and samples files, I think the 600 is equal or slightly better than Nikon version. 400 is visibly better than Nikon's 400 and adding 1.4x make it close to 600, but 600 is the way to go if that's the focal length you are after. I'm positive that 600 will beat 300+2X in IQ and AF....Show more →
There was some improvement to OSS with the A7RVI, but it still doesn't match what Nikon and Canon can do. This is one more reason why I hope Sony does a top end stacked APS-C sensor body. The smaller sensor will allow for better OSS and something around 40mp would be a great match with the 400mm f4.5.
Nice video by Jan. I'm looking forward to shooting with the lens, it will be interesting to see his next video with the lens comparisons. Personally, lacking in the OSS department doesn't impact me, I'm not a video shooter but it was a good catch on his part. I can see if you do shoot a lot of video this would be annoying to some degree. The little lens looks solid.
hoodlum90 wrote:
There was some improvement to OSS with the A7RVI, but it still doesn't match what Nikon and Canon can do. This is one more reason why I hope Sony does a top end stacked APS-C sensor body. The smaller sensor will allow for better OSS and something around 40mp would be a great match with the 400mm f4.5.
Something came to my mind when you mentioned APS-C body. Is the body size is the issue here? I don't know anything about the OSS technology. Sony is making camera and lens so small. Is that a limitation?
Jemini wrote:
Something came to my mind when you mentioned APS-C body. Is the body size is the issue here? I don't know anything about the OSS technology. Sony is making camera and lens so small. Is that a limitation?
It is not body size but rather stabilization technology. But an APS-C sensor is smaller and lighter than a FF sensor. That means the APSC sensor can be moved around more quickly with adjustments. It is also why m43 have always had very good stabilization.
hoodlum90 wrote:
It is not body size but rather stabilization technology. But an APS-C sensor is smaller and lighter than a FF sensor. That means the APSC sensor can be moved around more quickly with adjustments. It is also why m43 have always had very good stabilization.
Hope Sony will make a proper shaped APS-C camera with normal VF and controls. 6700 was a disaster to use with long lens. Z50II has perfect body (size, weight and shape) for APS-C. Just replace the sensor with a higher resolution stacked sensor (may be the rumored 26mp)
deevee wrote:
Another detailed review from Japan with comparisons to 300mm
Enjoy :-)
Interesting comparisons
400 f/4.5 vs 300mm + 1.4x : center sharpness looks the same. Edge sharpness looks better on the 300mm.
600 f/6.3 vs 300mm + 2x : both center and edge sharpness look extremely similar.
So it's basically a trade-off weight (and price) vs option to shoot at 300mm f/2.8 (+ 1/3 stop at 420 and 600)
Fboss wrote:
Interesting comparisons
400 f/4.5 vs 300mm + 1.4x : center sharpness looks the same. Edge sharpness looks better on the 300mm.
600 f/6.3 vs 300mm + 2x : both center and edge sharpness look extremely similar.
So it's basically a trade-off weight (and price) vs option to shoot at 300mm f/2.8 (+ 1/3 stop at 420 and 600)
I hope new lenses are better than 300 plus corresponding TC's. 300+2X is great for 2X combo. But not any miracle as a 600mm.
As far as the image stabilization goes. I think the 600mm’s super lightweight may be working against it somewhat
Usually, with a lens this long, the actual mass of the lens helps reduce small jerky movements of the hands and body working in tandem with the IS. I think the lightweight is allowing more of the natural body movements to be transmitted to the lens/body combo
That, and Sony not so great IS, is showing more with the longer focal length
It appears that testing was done like all other tests so far, at very small shooting distance.
Do you speak Japanese or see a clue to the actual distance to subject?
Waiting for someone to fill the frame with a Pelican or larger sized subject.
Fboss wrote:
Interesting comparisons
400 f/4.5 vs 300mm + 1.4x : center sharpness looks the same. Edge sharpness looks better on the 300mm.
600 f/6.3 vs 300mm + 2x : both center and edge sharpness look extremely similar.
So it's basically a trade-off weight (and price) vs option to shoot at 300mm f/2.8 (+ 1/3 stop at 420 and 600)
duncangr wrote:
I am pretty sure that's more or less how most of the modern mirrorless camera systems work in order to draw the subject detection boxes on screen and to determine the AF point(s), AF action required and AI based white balance adjustment.
And I am pretty sure that you missed the mandatory downsampling step because the bandwidth and power limitations of BIONZ does not allow it to operate the way you seem to think it does. You can't run inference at 120 fps on the full raw data stream under 4W, and that power budget is for everything, not just AI. The Nvidia TX2 runs face recognition in about 200ms on a 1920x1080 image. That's just 5 frames per second. While consuming 15 watts.
old-gregg wrote:
And I am pretty sure that you missed the mandatory downsampling step because the bandwidth and power limitations of BIONZ does not allow it to operate the way you seem to think it does. You can't run inference at 120 fps on the full raw data stream under 4W, and that power budget is for everything, not just AI. The Nvidia TX2 runs face recognition in about 200ms on a 1920x1080 image. That's just 5 frames per second. While consuming 15 watts.
I don't know first hand how Sony or Nvidia do things but on the iPhone there is no downsampling step required - you simply pass a reference to the video frame buffer - once you have set up the source i.e. live feed camera or video file, and the rectangle to track or the selected detected object to track (obtained when the user taps the screen to select the area or object) - and then you get back the tracked position in the frame and you draw or reposition your annotation on the display view.
TX2 is 2017, Sony A9 era, so probably not relevant. It made up of separate CPU, GPU and memory and not a unified architecture like the Apple silicon and latest Sony SOCs. Maybe that's what's in the Nikon Z8/Z9 hence the size and limited performance. - just kidding...