Modern photogrammetry has potential in many areas of human activity from construction to meteorology. Over the past five years, technologies and methods have been developed in this discipline that challenge traditional views and approaches in these fields. In particular, the introduction of structures on the methodology of movement has led to a significant increase in the use of photogrammetry in geological and engineering-geological practice. Achievements have been achieved mainly through social, political, environmental and technological changes around the world. Human mobility has increased significantly due to population growth, climate change and globalization. Innovations in photogrammetry have also been strongly influenced by the development of information and communication technologies, robotics and computer vision. With the use of remote sensing and radar techniques, the ability to collect, analyze and integrate data has greatly increased even for reflective surfaces such as metal or glass and uniform textures such as snow or ice. The availability of new high-resolution digital cameras and photogrammetry software has led to a gradual improvement in the quality of the surface data of the objects to be collected. Creating photogrammetric 3D models is now much faster and easier thanks to the use of motion structures but the use of ground-based control points to scale the model is still required. The paper considers the shortcomings of photogrammetry methods, quality problems and the need for additional processing of the final result. The reason for these shortcomings is that when using standard methods of photogrammetry, it is impossible for the software to evaluate the desired result and take any measures to improve it. After analyzing the general stages of the photogrammetry process, it is proposed to use computer vision and in particular the technique of object recognition using models based on machine or deep learning, which will help identify "places of increased attention" and remove useless data.
Looking at various articles or commercials, it may seem that today with the help of photogrammetry methods at home you can get a perfect digital representation of almost any object. This statement is true only in the case of expensive equipment and/or include post-editing processes that require special skills and additional software. Despite the fact that 3D scanning has been a great success in recent years and has become much more accessible, modern users have to be satisfied with mostly moderate results. In addition, even professionals often have a situation where after a long preparation of digital images, the end result is less detailed than expected.
The purpose of this work is as follows: to develop a way to improve the methods of photogrammetry so that the end result does not require additional post-processing.
In a broad sense, photogrammetry is defined as the use of a combination of hardware and software to analyze objects or environments from the real world, gather the necessary information and convert it into a digital model. Photogrammetry methods recognize position, shape and size based only on photographs.
This technology allows you to measure the object without the need for direct contact between the device and the object itself. Many approaches and techniques are used for this, from such disciplines as optics and projective geometry. One method of photogrammetry is to obtain three-dimensional measurements from information presented in two dimensions, for example, the problem of finding the distance between two points lying on a plane parallel to the plane of the photographic image can be solved by measuring the distance in the image if known scale. Another method allows you to implement a physically valid rendering, by obtaining from photographs of materials a range of colors and values representing albedo, metallicity, model shading and reflection.
Thanks to the common elements in digital images made from different angles and at different angles, computational algorithms are able to reconstruct the object.
The probability of possible errors in spatial orientation decreases if the angular distance between the lines of view is close to 90°. Moreover, when shooting, it is necessary to provide good lighting of the scene and each point must be clear on as many digital images as possible. Also, it is good practice to divide the space in front of the subject into equal parts (number of cameras plus one) to get the same angular distance between adjacent photos. In the case of working with three cameras, the angular distance is 180/4 = 45° and in the case of four cameras -180/5 = 36°. Then the angular distance between two non-adjacent chambers will be 90° in the first case, as shown in Figure 1 and 72° or 108° in the second case, respectively, as shown in Figure 2 [1].
Photogrammetric processing includes several stages that allow to obtain a two- or three-dimensional digital model of the object as a result of the process.
Not all photogrammetry projects are the same but almost every one of them includes the following stages:
Project planning. Includes a number of important considerations that affect the success of the final result. These include the selection of equipment and software, its settings, selection of place and time for shooting, planning the trajectory of the camera and the number of taken photos
Obtaining images and control information. There are a number of strategies for collecting digital images in the process of photogrammetry. Usually, the choice of strategy is guided by the chosen software or the type of processing used. Depending on this, a stereo or convergent image set may be required. Additional information about the external environment can be added to the project to adjust the placement of the model or to provide geometric constraints. If the photogrammetric model is to be placed within an existing frame of reference or unit of measurement (geodetic, cartographic or local), then sufficient external references defined in that frame must be integrated. Three-dimensional reference frame or units of measurement are determined by size, position and orientation. Typically, reference information is stored as checkpoints (points with known coordinates in the reference frame that identify the photo), the length of the objects that identify the photo or the angles between them. The minimum amount of information required to scale, orient and position a photogrammetric model is two three-dimensional and one one-dimensional control points. If the number of points provided exceeds the minimum value, then control information can be used to determine the shape of the photogrammetric model. In this case, it is necessary to make sure that the control information is at least 3 times more accurate than the photogrammetric model itself. If this is not the case, then redundant information will only distort the final model and will have a detrimental effect on its accuracy. It is also possible to apply control information after the three-dimensional model has been created. In this case, it will only affect the location of the model and will not cause distortion.

Figure 1: The first location, using 4 cameras

Figure 2: The second location using 4 cameras
Image processing and block triangulation. Most digital images taken in natural light require some digital processing, which may include adjusting the white balance, brightness, contrast or other properties of the image. It is important to note that in no case can you change the height or width of the image intended for photogrammetry. To obtain three-dimensional points from a two-dimensional image, it is necessary to triangulate at least two images. If more are used, then such a set is called a block. In order to triangulate the whole block, it is necessary to measure a sufficient number of control points. Constraints can also be placed on a specific set to provide angular, linear or planar properties. After successful triangulation, you can obtain and export a two-dimensional or three-dimensional model.
Creating and exporting the end result. A typical end result of a photogrammetry project can be two-dimensional vector graphics, a point cloud, three-dimensional polylines, raster graphics, a mesh object or surface. All of them must include the relevant metadata for each of the above stages of photogrammetry, as well as metadata for further processing and creation of the final file.
As mentioned more than once, the quality of the final model directly depends on the quality of digital images prepared for processing. Very often there will be situations when the parts of the model will not have enough detail, the necessary textures will lack clarity and the capture of useless data can not be avoided at all. Some shortcomings can be partially eliminated by post-processing, if you use specialized software or add new digital images to the unit. Additional manipulations increase the project implementation time and require special skills from the user.
All of the above problems have one thing in common-the software used to process a set of digital images knows nothing about the ultimate goal. For software, any object is a set of points with certain geometric data that must be connected. Unfortunately, it is not possible to obtain or transmit any information about what is being modeled and what deserves increased attention in a standard photogrammetry project. This information would greatly help to reduce project implementation time and improve the quality of the result.
Formed in a certain way information, is the result of each of the mentioned stages of photogrammetry. Its format and way to transfer take into account the next stage and processes that must be performed. Thus, the result of the second stage is a set of digital images with additional control information; the result of the third stage is a points cloud containing geometric information and texture properties; the result of the fourth - a two- or three-dimensional object with properties of a real object or surface. It can be quite difficult to manipulate the results of the third and fourth stages, considering complexity of their structure. A large number of points or landfills makes it almost impossible to adequately analyze. In addition, depending on the needs of the user, the result of the last stage may be completely different objects that can not be generalized. It should be noted that the result of the second stage is a set of digital images that have the same resolution, the same size and, at best, they all contain one common object. Undoubtedly, such information can help speed up the implementation of the photogrammetry project and improve the quality of the result.
To process a large number of digital photos and identify the depicted object, we 'll use computer vision. This discipline is a field of methods developing researches that help computers "see" and understand the content of digital images and videos. It shares the ideas and concepts with such disciplines as artificial intelligence, digital image processing, machine learning, deep learning, pattern recognition, probabilistic graphical models, scientific calculations and mathematics.
The purpose of computer vision is to understand the meaning of digital images. Object detection is a computer vision technique that allows you to distinguish and find objects in an image or video. Object detection is commonly confused with image recognition, it is important to find out the differences between them before proceeding. Image recognition assigns a label to the image. If there is a cat in the image - the image is labeled "cat". Object detection lable set of pixels, so each cat will receive its own label separately and its area in the image will be indicated by a separate rectangle. The method predicts where each object is located and which label should be used, as shown in Figure 3. Thus, object detection provides more information about the image than object recognition. Thanks to this type of identification and localization, it is possible to fully automate the stage of digital image processing in the photogrammetry project, eliminate unwanted elements and highlight the area that requires increased detail.

Figure 3: Example of Comparison of Image Recognition and Object Recognition
In general, object detection can be broken down into approaches based on either machine learning or deep learning. In more traditional machine-based approaches, computer vision techniques are used to highlight various features of an image, such as a color histogram, to identify groups of pixels that may belong to an object. These functions are then fed into a regression model, which assumes the location of the object along with its label. On the other hand, in deep learning approaches use a convolutional neural network to perform uninterrupted, uncontrolled detection of an object in which properties do not need to be determined and extracted separately.
Models of object detection based on deep learning usually consist of two parts. The encoder takes the image as input and passes it through a series of blocks and layers, which learn to extract the characteristics that will be used to search and mark objects. The source data of the encoder is then transmitted to the decoder, which predicts the bounding boxes and labels for each object.
The regressor is the simplest decoder. It uses the source data of the encoder and directly predicts the location and size of each bounding box. The initial data of the model are a pair of X, Y coordinates for each object and their dimensions in the image. This type of model, although simple, is limited - before processing, you must specify the number of probable objects. If there are two cats in the input image, one of them will go unnoticed if the model was designed to detect only one object. However, if the number of objects to be found on each digital image is known in advance, conventional regression models are a very suitable option.
The region proposal network is an extension of the regression approach. In this decoder, the model offers an image area where, in its opinion, the object can be located. The pixels of these regions are processed by the classification subnet to confirm or reject the label. The advantage of this method is that its model is more accurate and flexible and can offer any number of regions that contain objects. However, increasing accuracy reduces the speed of calculations.
Single shot detectors are looking for a compromise between the previous two approaches. Instead of using a subnet to offer regions, the DOZ relies on a set of predefined regions. Grid of anchor points, rectangles of different shapes and sizes, such as shown in Figure 4, is superimposed on the input image and serves to define the regions [2]. For each rectangle at each anchor point, the model predicts whether an object exists within a region or not and changes its location and size to match the characteristics of the object. Because anchor points can be located very close to each other and each of them contains several rectangles, the DOZ method usually produces many potential intersections. Therefore, post-processing should be applied to the original data of this method to discard most redundant predictions and select the best one.

Figure 4: Grid of Anchor Points
When using photogrammetry to obtain a digital model or surface, commonly user will face problems with the result object, such as: insufficient details in important areas; a large number of useless data are captured; These shortcomings can not be avoided with standard methods of photogrammetry, unless you spend extra time on post-processing using third-party software. I propose to use computer vision and object detecting, to pre-process digital images and determine the area of interest for triangulation. This approach can save time and money for the user and improve the quality of the final model.
Galantucci, L.M. et al.Accuracy Issues of Digital Photogrammetry for 3D Digitization of Industrial Products. National Polytechnic Institute of Grenoble, Italy, 2006, p. 13.
“Object Recognition Guide.” Fritz AI, Boston, 2017, https:// www.fritz.ai/object-detection/. Accessed January 2020.